Nitesh Singhal · Data & AI/ML Infrastructure

AI systems that hold up in production.

Most AI systems don't fail because the model is wrong. They fail because of a decision someone made about where a boundary goes, and it only shows up under real load. I've spent 12 years on the infrastructure side of that problem, and I help investors and engineering teams get those calls right.

  • Google
  • Uber
  • Microsoft
  • IIT Guwahati
  • Stanford GSB Executive Program

Three ways people bring me in. All of them start from the same place: a technical decision that costs real money to get wrong.

// diligence

Technical due diligence on data and ML infrastructure

You're about to write a large check into a company whose core claim is technical. I tell you whether the technology is real, whether it scales, and whether the team can execute.

  • Data platform and pipeline architecture, and the cost curve at 10x volume
  • ML infrastructure: serving, training, evaluation, feature stores
  • Inference economics, and whether the unit costs survive growth
  • Team and execution assessment against the roadmap they're selling
DeliverableWritten memo plus a partner call. Typical window is two to three weeks.
ForVC and PE investors, corp dev, founders preparing for diligence.
// review

Architecture and reliability review

Your generative AI feature works in a demo and falls over in production, or your inference bill is growing faster than your usage. Both are architecture problems, not model problems.

  • Reliability for generative AI: evals, guardrails, quality SLOs, failure modes
  • LLM inference cost structure, and where the money actually goes
  • Real-time and streaming pipeline design, and where it breaks under load
  • A second opinion on a design doc before you commit engineers to it
DeliverableA findings document with ranked risks and what to fix first.
ForSeries A through C engineering leaders, platform teams, CTOs.
// advisory

Standing technical advisor

Some decisions are hard to walk back, and you want someone on call who has already made them once. One standing call a month, plus async review when something big is in flight.

  • Architecture decisions with long half-lives, reviewed before they ship
  • Hiring calibration for senior and staff infrastructure engineers
  • Build versus buy on the AI platform layer
  • Somewhere to pressure-test a plan before it goes to your board
DeliverableMonthly retainer, capped hours, month to month.
ForFounders and heads of engineering at Series A/B companies.

Where I'm useful. Data platforms and pipelines, real-time and streaming systems, model serving, training and evaluation infrastructure, feature stores, and the cost and reliability of LLM inference. Computer vision is the model workload I know best.

Where I'm not. Robotics, actuation, and embodied systems. Pure research or model architecture work. Staff augmentation and implementation contracting. And nothing touching my employer's business, or that requires discussing it.

How this starts. A short call to check fit. I take a small number of engagements at a time, so I say no to most of them, and I'll tell you quickly if it isn't a fit.

12 yrs
data and ML infrastructure at Google, Uber, and Microsoft
Millions
of inferences per second, sub-second end to end
6
published essays on AI infrastructure

I work on the infrastructure underneath AI products: the data platforms that feed models, the serving layer that runs them, and the training and evaluation systems in between. The interesting part isn't the model. It's everything that has to be true for the model to be useful at scale, on a budget, without waking someone up at night.

Most of the problems I get pulled into don't start as code problems. They start as a decision someone made two years ago about where a boundary goes, and now the system can't do the thing the business needs. I like being in the room before that decision, and I'm useful in the room after it.

Google
Currently · ML infrastructure

Currently at Google, working on infrastructure for real-time machine learning: model serving, and the training and evaluation systems around it. Computer vision models are the workload I know best. Earlier, Google Assistant.

Uber
2018 to 2020 · Data platform

Built the real-time analytics platform that internal teams used to go from raw events to queryable data without standing up their own pipelines. Kafka, Flink, Pinot, Presto. This is where I learned what streaming systems do when they meet a real business.

Microsoft
2014 to 2016 · Dynamics PowerApps

Enterprise platform engineering, plus the analytics and monitoring architecture behind it.

Computer Science, IIT Guwahati. Stanford GSB Executive Program. ACM-ICPC regional semifinalist, 2012 and 2013.

Writing

I write about the same problems I get hired to look at.

Talks

Both on making generative AI reliable enough to ship.

DATA festival Online 2025

Reliability-first Generative AI: Turning Models into Business-Ready Systems

October 21, 2025

Generative models unlock powerful automation, but their tendency to produce convincing and incorrect output blocks adoption anywhere the stakes are real. A practical engineering approach for making generative AI auditable and reliable in enterprise workflows.

View event details
Summit of Things 2025

Hallucinations to High-Stakes Reliability: Building Trustworthy Generative AI Systems

October 21 to 23, 2025

How to build generative AI systems that hold up in production when the cost of a wrong answer is more than an awkward demo.

View event details

Have a technical decision you can't easily walk back?

Tell me the shape of it. If it's a fit I'll suggest a short call, and if it isn't I'll say so.

What's this about?