Expertise

Data platforms that turn raw inputs into usable intelligence

I work on data platforms as the layer that carries analytics from ingestion to delivery. That means orchestration, compute, metadata, warehouse layers, observability, and scaling controls working together as one production system.

What the work covers

The work spans ingest, transform, model, serve, observe, and scale. It connects orchestration, compute, metadata, and serving into one operating layer for analytical workloads.

  • Spark on Kubernetes and EKS-based compute orchestration
  • Dremio, dbt, Delta Lake, metadata, and warehouse layers
  • GitOps-managed infrastructure on AWS
  • Cost, performance, and reliability tuning for analytics systems

Where I’ve built this

TierraViva AI spans ingestion and mining pipelines, GitOps-managed AWS and Kubernetes infrastructure, analytical processing systems, serving layers such as APIs and web interfaces, and agent-ready biodiversity workflows.

Tools I use here

Each tool below has a specific role inside data platform engineering. Fullstack delivery tools (marked) are how this domain’s own UIs, dashboards, and APIs get built.

Tools

Apache Airflow pipeline orchestrationdbt data modeling and transformationApache Spark distributed compute and processingDelta Lake transactional storage and warehouse layerDremio serving/query engine (with KEDA-driven elastic autoscaling)Databricks ML and analytics computeGreat Expectations data quality validationPostgreSQL relational storeKEDA event-driven executor autoscalingAmazon EKS Kubernetes-managed analytics infrastructureArgoCD (GitOps) infrastructure-as-code delivery for the data stackFastAPI · React · TypeScript · WebSocket fullstack delivery (analytics APIs, dashboards, serving surfaces)

Related projects

data platform engineerbig data platformspark on kubernetesdremioairflow on eksanalytics infrastructure

Data

TierraViva AI

Biodiversity intelligence platform that ingests and mines policy, biodiversity, patent, and research data, processes it through GitOps-managed cloud data infrastructure, serves it through APIs and web interfaces, and makes it available to agents for analysis, navigation, reporting, and decision support. Its implementation spans document and research corpora, ETL and orchestration systems, Spark and analytics infrastructure, serving layers, and agent-consumable knowledge workflows. Serves the Data Platform Engineering and Applied Analytics domains.