AI Engineer

Yashwanth Reddy Boddireddy

Building Production RAG, LLM Applications & ML Infrastructure

RAG · LLM Evaluation · ML Serving · MLOps

94%

RAG Answer Relevance (Ragas)

40%

Hallucination Reduction

4

Featured AI Systems

Projects

Systems I Built

A showcase of my projects spanning production RAG systems, LLM observability, AI applications, and data analytics.

Nimbus Support Agent
Nimbus Support Agent
A LangGraph-orchestrated multi-agent support system, built phase by phase — from a no-persistence RAG bot to real Stripe refunds, real Shopify order lookups, real Zendesk ticketing, and a live ops dashboard. Guardrails like refund limits and read-only inventory access are enforced in code, not prompts, and every claimed outcome is independently verified against the real downstream API.
LangGraph
FastAPI
Chroma
Supabase
Next.js
Stripe
Shopify
Zendesk
Financial RAG Engine
Financial RAG Engine
A production-shaped RAG service for 10-K filings, built solo in 5 days and deployed to AWS ECS from day one. Hybrid dense (Pinecone) + sparse (BM25) retrieval fused with RRF, FlashRank reranking, and a Self-RAG grade-and-retry loop before answering with inline citations. Ragas-evaluated: 1.000 faithfulness, 1.000 context recall.
FastAPI
Pinecone
BM25
FlashRank
Self-RAG
Redis
Langfuse
Ragas
AWS ECS
LLM Monitoring & Observability Dashboard
LLM Monitoring & Observability Dashboard
Full observability layer over a production RAG system using Langfuse for tracing. Every query broken into retrieval, reranking, and generation time. Grafana dashboards track p50/p95 latency, token costs, and quality scores. Regression gating in CI blocks deploys when latency spikes or eval scores drop.
Langfuse
Prometheus
Grafana
Python
FastAPI
GitHub Actions
Real-Time AI Interview Assistant
Real-Time AI Interview Assistant
Next.js-based Real-Time AI Interview Assistant using GPT-4 with speech recognition and analytics dashboard for interview performance tracking.
Next.js
OpenAI GPT-4
Speech Recognition
Analytics
Career

Experience

From Electrical Engineer to full-stack data development, my path has been driven by a passion for solving complex problems with data and code.

May 2025

Master of Science in Data Science, Statistics @ New Jersey Institute of Technology (NJIT), Newark, NJ

Data Science
Statistics
Machine Learning
Deep Learning
Software Engineering
July 2025 - Present

AI Engineer @ Endeavour Technologies — Jersey City, NJ

  • Created an AI-first workflow using Claude Code to build internal tools: AI writes Python and React code, engineers review on GitHub before shipping, adopted by operations teams daily.
  • Built a financial document search system using Python, Pinecone, BM25, Cohere, LangChain, and OpenAI to retrieve and verify answers so analysts spend time on analysis instead of fact-checking.
  • Implemented query expansion and multi-hop reasoning to improve retrieval accuracy on complex financial queries, enabling the system to understand nuanced questions across multiple documents.
  • Implemented end-to-end LLM application tracing using Langfuse and Python to observe all application calls and identify bottlenecks, so engineers debug issues faster without guessing.
  • Built real-time dashboards using Prometheus, Grafana, MySQL, and Elasticsearch to track infrastructure and application health, enabling the team to catch and fix problems instantly.
  • Implemented end-to-end ML pipelines using Python and Kubernetes automating data ingestion, feature engineering, model training, packaging, and deployment with Ragas evaluation so models ship without manual bottlenecks.
  • Built feature pipelines with validation and versioning to ensure data quality across ML models, preventing model degradation from bad data entering production.
Claude Code
Python
React
Pinecone
BM25
Cohere
LangChain
OpenAI
Langfuse
Prometheus
Grafana
MySQL
Elasticsearch
Kubernetes
Ragas
July 2024 – September 2024

Software Engineer Fellow @ Headstarter AI — NYC, USA

  • Shipped production AI applications using Python, FastAPI, OpenAI, Pinecone, and Stripe handling real payments and customer data from design to deployment.
  • Implemented CI/CD pipelines using Python and GitHub Actions for automated testing, versioning, and reproducible deployments across dev, staging, and production environments.
  • Led engineers through full development cycles, reviewing architecture and code quality to ship projects on schedule without technical debt.
  • Deployed on AWS with GitHub Actions CI/CD and automated tests to catch problems before users encounter them, ensuring system reliability as requirements changed.
Python
FastAPI
OpenAI
Pinecone
Stripe
GitHub Actions
AWS
January 2023 – August 2023

Software Engineer @ Google (client via Accenture) — Hyderabad, India

  • Fixed Bard's (now called Gemini) harmful responses using Python regex patterns to filter political bias, sexual content, hate speech, and misinformation so users got helpful answers.
  • Validated content filtering by regenerating responses in production before releases, confirming harmful content was blocked.
Python
Content Moderation
Responsible AI
January 2022 – September 2023

Software Engineer @ Accenture — Hyderabad, India

  • Built NLP classification and ranking models using Python and transformer-based architectures to replace generic recommendations, deployed across apps so users found relevant content.
  • Automated ML workflows using Python automating data prep, model training, testing, packaging, and deployment to eliminate manual bottlenecks in release cycles.
  • Shipped full-stack features across .NET, C#, React, Angular, and Node.js for enterprise clients, owning the entire path from API design to production.
  • Optimized REST endpoints and SQL Server databases using Python and SQL, tuning queries and adding caching so users got instant responses.
  • Mentored junior engineers on building and shipping ML systems; two led independent AI projects for clients.
Python
Transformers
NLP
.NET
C#
React
Angular
Node.js
SQL Server
Mentorship
Skills

Stack

The tools I build production AI systems with, ordered by where my depth actually sits.

AI & LLM Systems

Model frameworks, retrieval, and LLM tooling.

PyTorch

LangChain

OpenAI

Hugging Face

TensorFlow

Scikit-learn

MLflow

MLflow

Databricks

ML Infrastructure & DevOps

Deployment, orchestration, and cloud infrastructure.

Kubernetes

Docker

AWS

Azure

Google Cloud

Apache Airflow

Jenkins

Kafka

GitHub Actions

Linux

Backend & APIs

Services and datastores behind the AI systems.

FastAPI

Flask

PostgreSQL

MySQL

MongoDB

Data & Analytics

Processing, analysis, and reporting.

Python

Pandas

NumPy

Apache Spark

Hadoop

Snowflake

dbt

R

Tableau

Power BI

Jupyter

Matplotlib

Matplotlib

Frontend

Used for internal tools and project frontends.

JavaScript

TypeScript

React

About Me

About

Yashwanth Reddy Boddireddy

Yashwanth Reddy Boddireddy

AI Engineer specializing in production GenAI and ML systems, helping organizations leverage data for competitive advantage. With experience at Endeavour Technologies, Headstarter AI, and Accenture (including a client engagement with Google), I focus on developing intelligent systems that enhance user experiences.

I combine technical expertise with business acumen, following a systematic approach: understanding requirements, designing data-driven architectures, and implementing scalable solutions with measurable results.

Contact

Let's Talk

Open to full-time AI Engineer roles. If your team is hiring or you'd like to talk about a fit, I'd love to hear from you.

Contact Information
Feel free to reach out through any of these channels

Hiring for an AI Engineer role?

Connect with me

Send a Message
Fill out the form below and I'll get back to you as soon as possible