(415) 555-0149 ◇ San Francisco, CA
ishaan.mehta@example.com ◇ linkedin.com/in/ishaan-mehta ◇ ishaanmehta.com
LLM engineer with 5 years building production systems on large language models, currently the lead LLM engineer for a document intelligence product processing 40 million pages a month for 600 enterprise customers. Built the retrieval-augmented generation platform that raised answer accuracy from 71% to 91% on a 5,000-question benchmark, cut inference cost per query 64% through model routing, caching, and fine-tuning, and holds a 99.95% availability record on the inference service over 2 years. Fine-tuned 6 models that beat the general-purpose baseline by an average of 14 points on domain tasks. Builds language model systems that are measured, monitored, and cheap enough to run at scale.
Thesis on efficient fine-tuning of transformer models. Coursework in Deep Learning, Natural Language Processing, Distributed Systems, and Information Retrieval.
Graduated with distinction. Coursework in Algorithms, Machine Learning, Databases, and Operating Systems. Also holds an AWS Certified Machine Learning Specialty certification earned in 2022.
- Lead LLM engineering for a document intelligence product processing 40 million pages a month for 600 enterprise customers, leading a team of 7 engineers.
- Built the retrieval-augmented generation platform with hybrid search, a reranker, and structured context assembly that raised answer accuracy from 71% to 91% on a 5,000-question benchmark and cut unsupported answers 80%.
- Cut inference cost per query 64% through model routing across 3 model tiers, semantic caching with a 38% hit rate, and 6 fine-tuned models, and hold a 99.95% availability record on the inference service over 2 years at 20 million queries a month.
- Built semantic search and question answering systems for an enterprise search product used by 200 customers over 2.5 years.
- Shipped the first language model based answer feature in 2022, reaching 78% accuracy on a 2,000-question benchmark and a 30% click-through lift.
- Built a training pipeline that fine-tuned embedding models on 4 million query pairs, raising retrieval recall at 10 from 62% to 84%.
- Built a passage reranking model during a 14-week internship that raised top-1 accuracy 18% on an internal benchmark.
- Wrote an evaluation harness for 3 retrieval models used by the team for 2 years.
- Received a full-time offer at the end of the 14-week internship based on the reranking model.
Retrieval-Augmented Generation Platform. Designed a platform with layout-aware chunking for 40 million pages a month, hybrid dense and keyword retrieval over 2 billion chunks, a fine-tuned cross-encoder reranker, structured context assembly with citation tracking, and an evaluation harness of 5,000 questions across 12 document types, which raised answer accuracy from 71% to 91%, cut unsupported answers 80%, and serves 20 million queries a month.
Inference Cost Programme. Built a model routing layer classifying queries into 3 complexity tiers with a fine-tuned 7-billion parameter model handling 55% of traffic, semantic caching with a 38% hit rate, continuous batching on vLLM, and quantisation for the smaller tiers, validated against the full benchmark, which cut cost per query 64%, saving about $4.8M a year, with a 1-point accuracy change.
Domain Fine-Tuning Programme. Fine-tuned 6 open-weight models with parameter-efficient methods on 400,000 curated examples across extraction, classification, and summarisation tasks, with a data pipeline for labelling, deduplication, and quality filtering, which beat the general-purpose baseline by an average of 14 points on domain benchmarks and cut latency 60% on those tasks.
- Speaker at 4 machine learning conferences on retrieval-augmented generation and inference cost.
- Maintain an open-source evaluation library for retrieval systems with about 2,500 stars on GitHub.
- Volunteer mentor for a machine learning fellowship, coaching about 4 fellows a year.
- Lead an LLM engineering team of 7 with design reviews, roadmap ownership, and hiring, with 2 promotions in 2 years.
- Own the inference platform budget of $7M a year and present cost, accuracy, and reliability metrics to the chief technology officer monthly.
- Wrote the company standard for language model evaluation and release adopted by 4 product teams.

.webp)



.webp)


%20Which%20Should%20You%20Use.webp)