Introduced SRE practices (SLI, SLA, SLO, Error Budget) to the Data Platform, root-causing issues and driving fixes
Ingested and mapped data from logs and system tables provided by Databricks, among other sources
Built dashboards used by both engineers and management
Optimized hundreds of data pipelines by profiling bottlenecks via Spark UI, reducing compute cost
In the age of Generative AI, built an LLM agent that evaluates Change Requests against SRE standards, catching errors before they reach the platform
Precompute facts in Python (AST-based test-import mapping, git diff scope, coverage %) and inject them as fixed values into the LLM prompt, eliminating hallucination and score drift and making evaluations deterministic and reproducible across sessions
Replaced a single 600-line/40-anti-pattern prompt with 3 narrowly-scoped LLM calls (7 critical patterns always, 20 quality / 13 technical patterns conditionally), fixing inconsistent detection at scale
Built two-tier git caching (shared bare mirrors + per-task worktrees) to parallelize repo access across concurrent evaluations without lock conflicts
Senior Data Engineer
Built Data Governance tooling
Data Veracity tool ensuring data reliability via metrics like volume and velocity, from source systems to the insight layer, applying Statistical Process Control (learned from historical data) to enforce quality standards
Data Download Tool governed by the Data Governance team, enabling centralized requests via Collibra with input-query safeguards to block dangerous actions
Owned infrastructure for the Kappa data system team
Built an IaC framework for deploying pipelines
Integrated the monitoring system
Jan 2022 - Oct 2022
-
Cake by VPBank
-
Data Engineer
Key contributions:
Built a data transfer pipeline
Retrieved data from SFTP servers and Oracle Object Storage
Parsed and transformed structured text, XML, and CSV data
Refactored and improved the in-house data framework codebase
Planned platform monitoring and data lineage collection/analysis with the team
May 2021 - Dec 2021
-
BE GROUP JSC
-
Data Engineer
Supported the team in:
Developed new Airflow pipelines for MongoDB, ClickHouse, Salesforce, Google Cloud Pub/Sub, and reporting data sources
Maintained the in-house data framework (data pipelines, CDP)
Jul 2019 - May 2021
-
Galaxy Play JSC
(formerly
Fimplus JSC
) -
System Developer
Built an in-house CDN cache provider using Adaptive Replacement Cache (ARC) and Consistent Hashing; RPC server in Go, ARC module for Redis in C
Built a watch-rights Go service: device-aware video manifest matching (graph traversal), rule evaluation, secure play-link generation, and session management
Redis as the database (server-side sentinel + client-side ring)
RabbitMQ for cross-service events; Kafka tracker consuming from an event store to maintain session validity
Built a multi-profile module for the account management service
Maintained legacy systems
Mar 2018 - May 2019
-
AlphaSC Co Ltd
(formerly
Agilsun Co Ltd
) -
Software Engineer
Jan 2016 - Jan 2018
-
Apparent Labs Pte Ltd
-
Software Engineer