Background & Professional Journey

With over 14 years of industry experience, I specialize in designing, building, and scaling high-throughput data platforms and ETL/ELT pipelines that drive business-critical analytics and experimentation at scale. Currently operating as a Lead Data Engineer, I bridge the gap between core data infrastructure and artificial intelligence, developing production-grade data pipelines alongside practical LLM training and modeling curricula.

Throughout my career, I have focused on architecting end-to-end data ecosystems from ingestion and stream processing to storage, analytics, and self-serve AI enablement across organizations.

Why AI from Scratch?

Recently, I decided to look inside the black box of machine learning. Instead of merely calling high-level API wrappers, I am trying my hand at building language models from the ground up: coding tokenizers, writing custom attention loops, and running local transformer validation on consumer GPUs.

Building from scratch builds invaluable technical intuition. When a production data pipeline breaks or a tensor dimension mismatch appears, having a foundational understanding of the underlying math and distributed systems architecture separates true engineering rigor from guesswork.

Engineering Philosophy

"Great data platforms enable great teams. Reliability and clarity always win over raw complexity."
Engineering Principle Legacy Approach Our Scalable Standard
Reliability > Raw Speed Brute-force speed with silent failures Trustworthy data with schema validation
Simple Architectures Over-engineered 10-tier microservices Clean, modular pipelines with low MTTR
Automated Data Quality Manual debugging after downstream alerts Strict ingestion validation & anomaly alerts
Full Observability Black-box batch jobs with no logs End-to-end lineage & latency tracking

Current Focus & Research

  • LLM Architecture & Fine-Tuning: Developing comprehensive curriculum materials and detailed code implementations for transformer training, GPT-2 validation, and domain-specific fine-tuning.
  • Surprisal & Linguistic Modeling: Engineering professional, production-grade codebases for psycholinguistic predictability and surprisal modeling.
  • Modern Lakehouse Architectures: Designing high-concurrency, cost-optimized data lakehouse implementations.

Technical Skills & Infrastructure

Cloud & Infra
AWS (EMR, Glue, Lambda, S3), Google Cloud (BigQuery), Docker, Kubernetes, Terraform
Big Data & Processing
Apache Spark (PySpark), Apache Kafka, Apache Airflow, Databricks, Hive, Presto
Databases & Storage
Snowflake, PostgreSQL, Elasticsearch, Redis, MongoDB, DynamoDB, Parquet/ORC
Languages & AI / ML
Python, SQL (Advanced), Java, Bash, PyTorch, Hugging Face Transformers

Get in Touch

Open to discussing data platform architecture, high-scale ETL, and LLM engineering. Feel free to reach out to me on LinkedIn or Twitter / X!