Kevin Hernandez

Kevin Hernandez

Senior Data Engineer, AI & Data Solutions

Designing scalable data pipelines, lakehouse architectures, and real-time processing across AWS, GCP, and Azure. Expert in Apache Spark, Databricks, Delta Lake, Kafka, and AI/ML integration.

Scroll
— About

Senior Data Engineer with extensive experience designing scalable data pipelines, lakehouse architectures, and real-time processing across AWS, GCP, and Azure.

Led the delivery of a HIPAA-compliant data platform for 500+ clinicians, achieving 93% extraction accuracy and reducing manual effort, and built a high-throughput image generation pipeline supporting 5,000+ requests per month.

Expert in Apache Spark, Databricks, Delta Lake, Kafka, and AI/ML integration (RAG, LLMs), with strong data-modeling skills for analytics and BI.

My goal is to leverage this expertise to drive data-driven solutions that improve operational efficiency and support strategic decision-making.

— Experience

Key Responsibilities & Achievements

  • Designed and built the backend data architecture for ZOYO, an AI powered real estate platform offering 8+ tools including virtual staging, interior/exterior design, day to night conversion, and image enhancement, scaling to 100+ monthly active users.
  • Developed a modular NestJS backend with domain‑driven data models, REST and WebSocket APIs, and JWT authentication, delivering all platform services with 99.9% uptime on AWS ECS with auto‑scaling.
  • Built a high throughput image generation and processing pipeline using RunPod and Flux models, coordinating data flow across storage, queues, and inference services to handle 5,000+ image requests per month at an average of 6 seconds per response.
  • Engineered data integrations with GPT 5 and Claude Sonnet 4.5 to power property description generation and context aware design recommendations, improving user engagement metrics by ~38%.
  • Implemented an Angular frontend with NgRx state management and RxJS reactive data streams for optimistic UI updates on image generation workflows, reducing perceived load times by ~45%.
  • Integrated Stripe billing data and built an ROI calculator to support pricing decisions, contributing to a 22% increase in free to paid conversion within 3 months of launch.

Technologies Used

NestJSAWS ECSRunPodFluxGPT-5ClaudeAngularNgRxStripeWebSockets

Key Responsibilities & Achievements

  • Led the design and launch of a HIPAA-compliant AWS data platform (EKS, Bedrock, SageMaker) with RBAC and audit logging, enabling real-time voice, text, and document data pipelines for 500+ clinicians and care coordinators in the first year.
  • Engineered a low latency document processing pipeline combining AWS Textract with Amazon Bedrock and RAG based EHR context injection, sustaining over 93% extraction accuracy and increasing downstream clinical report completeness by ~28%.
  • Built a lakehouse architecture using Databricks and Delta Lake to unify structured and unstructured clinical data, and applied Apache Spark (PySpark, Spark SQL) with star schema data modeling to process high volume datasets for downstream analytics and BI reporting.
  • Built automated data orchestration systems with MCP‑compliant interfaces using n8n and GitHub Actions for care‑plan routing, escalation handling, and follow‑up automation, cutting manual data‑entry errors by 30% and improving clinical operational efficiency.
  • Scaled FastAPI microservices on EKS with containerized inference workloads via Amazon ECR, enforcing HIPAA aligned network policies and PHI data isolation, achieving 99.8% uptime while reducing cloud spend by ~20% through autoscaling.
  • Implemented continuous data quality and LLM evaluation pipelines with RAG relevance scoring, hallucination detection, and PHI redaction monitoring, enabling proactive regression detection across production environments.

Technologies Used

AWSEKSBedrockSageMakerDatabricksDelta LakeApache SparkTextractRAGMCPn8nFastAPI

Key Responsibilities & Achievements

  • Containerized and deployed DeepHow's platform on GCP (GKE, Cloud Run, Cloud Storage) using Docker and GitHub Actions CI/CD pipelines, improving deployment reliability and reducing infrastructure costs by ~20%.
  • Built internal data pipelines and star schema data models using Apache Spark (PySpark, Spark SQL) and FastAPI, enabling clients to track workforce knowledge usage through Next.js dashboards and cutting analytics turnaround from days to hours.
  • Built Power BI dashboards and semantic models using DAX and Power Query/M to visualize workforce training and knowledge usage data, giving clients self-service analytics on top of the underlying data pipelines.
  • Developed core FastAPI backend services and GraphQL APIs powering DeepHow's AI knowledge engine, improving knowledge transfer workflows by ~30% and reducing customer support tickets by ~18%.
  • Built and maintained a customer-facing data application using Next.js with TypeScript and GraphQL, serving 50,000+ monthly users across manufacturing and industrial clients.
  • Mentored junior engineers on data pipeline design, TypeScript, and Docker best practices, and partnered with product managers to improve sprint planning and reduce post-release defects by ~25%.

Technologies Used

Apache SparkPower BIGCPGKEFastAPIGraphQLNext.jsTypeScriptDockerGitHub Actions

Key Responsibilities & Achievements

  • Designed and optimized PostgreSQL database schemas using conceptual, logical, and physical data modeling techniques, including star schema design, for IoT device and real estate data, improving query performance by ~35%.
  • Integrated Kafka to stream real-time IoT sensor data from smart buildings, enabling live tracking of energy usage, occupancy, and facility operations while reducing data latency by ~40%.
  • Built backend services and REST APIs using Django and Django REST Framework to support smart building and real estate portfolio management, accelerating feature delivery by ~25%.
  • Developed APIs with Firebase authentication, validation, and error handling, improving API reliability and reducing error incidents by ~20%.
  • Deployed backend services on Azure (App Service, Functions, Event Hubs) using Docker, and built Azure Event Hub based streaming pipelines to complement Kafka for real-time data ingestion, achieving 99.9% uptime and reducing infrastructure costs.
  • Automated internal facility and portfolio reporting workflows using Power Automate and Power Apps with custom connectors, reducing manual effort for smart building operations teams.

Technologies Used

PostgreSQLKafkaData ModelingDjangoAzureEvent HubsPower AutomatePower AppsFirebaseDocker

Key Responsibilities & Achievements

  • Built a full stack HR and recruiting platform, developing a React single-page application with MUI components on the frontend and FastAPI services backed by MongoDB on the backend, accelerating candidate processing and reducing onboarding time for outsourcing clients.
  • Delivered end-to-end product features across the full stack, including candidate tracking, job posting management, and employee onboarding interfaces, reducing manual HR workload for clients.
  • Participated in Agile ceremonies, code reviews, and CI workflows with Docker across frontend and backend codebases, improving sprint velocity by ~15%.

Technologies Used

ReactMUIFastAPIMongoDBDockerAgile
— Education

Achievements

  • Specialized in Data Engineering and AI
  • Core focus on Algorithms and Distributed Systems

Activities & Leadership

  • AI Research Group Member
  • Open Source Contributor

Relevant Coursework

Data Structures & AlgorithmsArtificial IntelligenceSoftware EngineeringDatabase SystemsCloud ComputingMachine Learning
— Skills

Technical Skills

Data Engineering & Pipelines

Apache Spark
Databricks
Delta Lake
Lakehouse
Kafka
Azure Event Hub
AWS Lambda
Step Functions
ETL/ELT
RAG Pipelines
Streaming

Data Modeling

Star Schema
Dimensional Modeling
Conceptual/Logical/Physical
Data Warehousing

Power Platform & BI

Power BI (DAX, Power Query/M)
Semantic Modeling
Power Apps
Power Automate
Custom Connectors

Cloud & DevOps

AWS
GCP
Azure
Docker
GitHub Actions
Jenkins
Vercel
CI/CD

Databases

PostgreSQL
MySQL
MongoDB
DynamoDB
Firebase
Supabase
Redis

Data Science & AI/ML

OpenAI
Claude
Gemini
DeepSeek
LangChain
LangGraph
LangSmith
RAG
MCP
Multimodal AI
Multi-Agent AI

Languages

Python
JavaScript (ES6+)
TypeScript
Go
Java
C#

Backend

Node.js
ExpressJS
NestJS
Django
Flask
FastAPI
REST
GraphQL
Laravel

Frontend

React
Vue.js
Next.js
Angular
Tailwind
MUI
SCSS

Workflow Automation

Vellum AI
Zapier
Make
n8n
Pipedream

AI Coding Assistants

Claude Code
Cursor AI
Google Antigravity
GitHub Copilot

AI No/Low-Code Platforms

Lovable
Bolt.new
v0
Replit Agent
Builder
Base44
Bubble
— Certifications

Professional Certifications

Industry-recognized credentials demonstrating expertise in AI, cloud computing, and modern technologies.

Generative AI Essentials on AWS
AI & Machine Learning

Generative AI Essentials on AWS

AIS Academy

AI in Finance Micro-Credential Certificate
AI & Machine Learning

AI in Finance Micro-Credential Certificate

IMA | Institute of Management Accountants

Generative & Agentic AI
AI & Machine Learning

Generative & Agentic AI

IBM Consulting

Generative & Agentic AI Expert Architect
AI & Machine Learning

Generative & Agentic AI Expert Architect

IBM Consulting

Azure Stack HCI Foundation Level
Cloud & DevOps

Azure Stack HCI Foundation Level

Dell & Microsoft

NABA Artificial Intelligence Certified
AI & Machine Learning

NABA Artificial Intelligence Certified

NABA Inc.

— Services

Project Catalog

Strategic consulting packages designed to accelerate your data platform, analytics, and AI initiatives with proven methodologies.

Data Pipeline Architecture Audit
ETL/ELTArchitecture

Data Pipeline Architecture Audit

ETL/ELT Assessment & Modernization Roadmap

12 days
View details
Lakehouse Architecture Design
DatabricksDelta Lake

Lakehouse Architecture Design

Databricks, Delta Lake & Spark Platform

14 days
View details
Real-Time Streaming Pipeline
KafkaStreaming

Real-Time Streaming Pipeline

Kafka & Event-Driven Data Architecture

14 days
View details
Power BI Analytics Platform
Power BIDAX

Power BI Analytics Platform

Semantic Models, DAX & Self-Service BI

10 days
View details
Data Warehouse & Dimensional Modeling
Star SchemaData Modeling

Data Warehouse & Dimensional Modeling

Star Schema Design for Analytics at Scale

10 days
View details
RAG Data Platform Blueprint
RAGLLM

RAG Data Platform Blueprint

Retrieval Pipelines for LLM Applications

12 days
View details
LLM Data Integration Strategy
LLMAPI

LLM Data Integration Strategy

GPT, Claude & Gemini in Your Data Stack

8 days
View details
MCP Data Orchestration Plan
MCPOrchestration

MCP Data Orchestration Plan

Model Context Protocol for Data & AI Workflows

12 days
View details
AI Media Processing Pipeline
InferenceGPU

AI Media Processing Pipeline

High-Throughput Image & Inference Workloads

14 days
View details
— Testimonials

Client Reviews

Trusted by CTOs, Engineering Leaders, and Founders to deliver enterprise-grade data platforms and AI solutions

Mark Stephens

Mark Stephens

CRO

DeepHow

"Kevin transformed our analytics capability. The data pipelines and Power BI dashboards he built took our reporting turnaround from days to hours, and our clients finally got self-service insight into their workforce training data. His work directly strengthened our customer retention."

Project: Analytics Data Pipelines
10+
Years Experience
500+
Clinicians on Platforms Built
99.9%
Production Uptime
93%+
RAG Extraction Accuracy
— Contact

Let's work together

If you would like to discuss a project or just say hi, I'm always down to chat.

© 2026 Kevin Hernandez. All rights reserved.