Jake Williams

Data & AI Platform Engineer — Kafka CDC Streaming · Snowflake Lakehouse · Agentic Analytics (Claude, MCP)

Denver, CO · info@jakeawilliams.com
linkedin.com/in/jakeintech · github.com/Jakeintech
jakeawilliams.com

Platform engineer with end-to-end ownership of an insurer's data platform: Kafka CDC streaming on Kubernetes, a governed Snowflake medallion lakehouse, and a production GenAI semantic layer. Spearheaded the enterprise Claude plugin marketplace — 12 plugins across 6 business domains, with ~200 monthly users, analyst to C-suite — and cut $30K+/yr in platform cost. Agentic analytics in production with MCP, model evaluations, and an audit trail; answers that hold up under an executive, an auditor, and a regulator.

Experience

Vantage Risk — Data Engineer III, Platform & Semantic Analytics

Mar 2024 — present
  • Spearheaded the enterprise Claude plugin marketplace (12 plugins, 6 domains): ~200 monthly active users — 40% of the company, incl. 5 C-suite regulars — 2K+ queries/month; cut token consumption 30% vs. unassisted LLM use; own roadmap, evaluations, and governance.
  • Own the enterprise CDC/replication platform (real-time event streaming): inherited mid-build, shipped it, re-architected Event Hubs → Apache Kafka (Strimzi on Kubernetes), cutting $30K+/yr; Debezium CDC from SQL Server, PostgreSQL, and Oracle (LogMiner), plus custom Java Kafka Connect connectors where none existed.
  • Lead engineer, semantic analytics (Claude + Snowflake Cortex): governed NL analytics with consistent regulatory reporting — Quarterly Reserve Review 3 days → <30 min, QBR prep 4 days → <1 hr; Cortex evaluations made answers trusted and compliant.
  • Analytics engineering at scale: 700+ dbt models (ELT) across 5 domains on a Snowflake medallion lakehouse, incl. a financial-reconciliation framework; type-2 period-fact models for point-in-time analysis; data-quality tests in CI.
  • Expert-level data orchestration (Dagster) — software-defined assets, partitioned assets, jobs, schedules, asset checks; run-failure sensors with email alerting for rapid pipeline-failure detection; business-driven sensors on bound deals improved responsiveness and closure 20%.
  • GitOps everything: Argo CD + Helm deliver every platform service — Kafka, connectors, orchestration, observability; PySpark compaction jobs on Apache Iceberg keep streaming tables query-efficient.
  • Observability as code: 60+ dashboards/monitors ship via GitHub Actions (Datadog, Prometheus); forecast monitors prevented 5 OOM incidents and avoided ~20% infra cost (FinOps).
  • Technical leadership: code reviews, mentoring, office hours driving AI-assisted development adoption; SOC 2 SOPs and audit support.

Vantage Risk — Reporting Analyst & Architect

Mar 2023 — Mar 2024
  • Co-designed the enterprise data architecture (lake, EDW, reporting layer) on the core design committee; helped build the reporting solution the company scaled on. Promoted to Data Engineer in 12 months.
  • Built a key component of the enterprise pricing tool — a TypeScript microservice on Azure integrating external and internal storage/processing APIs (GraphQL, MCP, PostgreSQL); SME on feasibility, cost, and risk.

State Farm — Data Analyst

Jan 2021 — Mar 2023
  • Drove data-governance initiatives for specialty P&C analytics — integrity, lineage, regulatory compliance; delivered an Agile migration of analytics workloads to AWS Redshift; designed and ran the team's AWS training program.

Independent Contractor — data, analytics & technology

2016 — 2021
  • Independent contractor building client data, analytics, and technology platforms — flagship: Laughlin River Tours' full stack (automated pipelines, AWS QuickSight over Glue); payment processing (Stripe) and menu systems for Laughlin Ice @ Regency Casino; Arizona Snowbowl table-management app ($100K+ cost avoidance); WrapWorks Las Vegas and others.