Ankit Raheja on Building Production AI Systems That Scale Beyond the Pilot Phase

Facebook
X
WhatsApp
Table of Contents
Ankit Raheja on Building Production AI Systems That Scale Beyond the Pilot Phase

You’ve built AI systems across three very different contexts: government fraud detection, enterprise automotive software, and Fortune 10 retail. What’s the biggest gap between how AI evaluation is discussed in industry conversations versus how it actually needs to work when systems go into production?

Industry conversations optimize for business value, which comes from aggregate accuracy. Production systems need to ensure the edge case scenarios are handled for high-stakes instances. Trust is built over time but can be lost with one big mistake. In fraud detection at SCIF, a 95% accurate model that misses the one organized ring costs you the program. For ranking models for Walmart, we need to optimize for not only short-term GMV gains but also long-term customer lifetime value. At CDK with AIVA, the hardest part was not building the eval set, it was getting sales, legal, and dealer ops to agree on what “good” even meant. In the end, evaluations need to be not just focused on average but also on the failures that can’t be ignored. Hence, need for a comprehensive Eval set with guardrails metrics in place.

You diagnosed and relaunched a twice-failed Identity Graph product at CDK Global. When you’re brought in to fix a failing ML initiative, what are you looking for first, is it usually a data problem, a product strategy problem, or something organizational?

I look for the unspoken disagreement about what “success” means. Twice-failed initiatives almost always have stakeholders who never aligned on the core question. At CDK, Identity Graph had been pitched differently to engineering (an infrastructure investment), to sales (a revenue lever), and to executives (a competitive moat). Each constituency was measuring success against a different yardstick, so any outcome looked like failure to someone. That’s not a data problem. That’s a problem of nobody having forced the conversation, and it was not aligned about who would be owning it as the data sources and data sinks can be different system and it requires a STO (Singled Threaded Owner).

You built 80+ enterprise APIs deployed across thousands of automotive dealerships processing billions in transactions. What did that teach you about designing for enterprise adoption at scale, and how does API platform thinking translate to building ML systems that need similar reliability?

At CDK, the lesson wasn’t really about APIs.  It was that enterprise adoption is won or lost on predictability, not capability. Dealers didn’t care that an endpoint was elegant; they cared that it returned the same shape, the same latency, the same error semantics every time, because their F&I workflows and DMS integrations couldn’t absorb surprises. That meant versioning was a contract, deprecation was a multi-quarter negotiation, and observability had to be good enough that I could tell a dealer principal why their deal jacket failed before they called support. ML System thinking is similar. On BuyBox, the work that actually moves adoption isn’t the model lift; it’s the shadow evaluation harness, the guardrails on score drift, the rollback story when a signal source degrades, and the explainability layer that lets a category manager trust why a seller won the Buybox.

You’ve shipped both GenAI platforms and supervised learning systems at production scale. When you’re scoping a new product problem, what determines whether GenAI or traditional ML is the right approach? Where does each actually win?

The real answer is use-case specific, and the framing of GenAI versus traditional ML is usually a false choice. My filter is three things: cost-per-decision economics, tolerance for variance, and whether the problem has a stable label. Supervised ML wins on high-volume, low-latency, well-labeled decisioning where you can A/B test and attribute lift. BuyBox ranking is the canonical case; a 40ms p99 budget makes an LLM call structurally impossible. GenAI wins on long-tail problems where labeling cost exceeds inference cost and natural language is the interface. AIVA at CDK fit because dealer service questions had infinite intent surface and no clean label set. The interesting design space is the hybrid: LLM as the reasoning and routing layer on top of supervised models doing the heavy lifting. Let GenAI own the messy human interface, let traditional ML own the deterministic high-volume decisioning underneath.

Your GenAI platform at CDK Global achieved $2M ARR serving car dealerships, while your ML ranking systems now serve hundreds of millions of consumers. What do those two experiences reveal about when conversational AI makes sense versus when recommendation infrastructure delivers better outcomes?

The two experiences map cleanly to a deeper distinction: conversational AI wins when the user knows what they want but cannot or busy to express it in a structured query, recommendation infrastructure wins when the user does not explicitly know what they want and the system has to infer it from behavior. AIVA worked at CDK because dealer service advisors had specific intent (“how many cars did i sell”) but no patience for navigating menus, so natural language collapsed a ten-click workflow into one turn. BuyBox works at Walmart because the shopper typing “running shoes” has not actually told you anything useful, and the lift comes from ranking against revealed preference at scale. The economics follow the same split. Conversational AI monetizes through workflow compression and seat- or usage-based pricing, which is why $2M ARR across a few thousand dealerships pencils out. It also depends on the size of the market and willingness to pay. Recommendation infrastructure at Walmart scale monetizes on transaction volume and small percentage lifts on a massive base, which is why basis points of GMV matter more than feature richness. The mistake I see often is teams reaching for chat interfaces on problems where the user has no articulated intent or building recommender systems for workflows where the user could have just told you what they wanted.

You started building production AI systems in 2013 with fraud detection for California’s workers’ compensation fund. What promises from that era of machine learning actually held up, and what got quietly abandoned? What does that pattern teach you about evaluating today’s AI claims?

2013 was a useful vantage point because the field was making two kinds of promises, and only one aged well. What held up is the boring stuff. Feature engineering as the real source of lift, clean labeled data beating algorithmic novelty, ensembles as the workhorse, and the truth that most production value lives in the operational wrapper around the model. At SCIF we caught fraud because we spent months with claims adjusters understanding what suspicious billing actually looked like, not because gradient boosting was magic.

What got quietly abandoned is more telling. Self-improving models that learn continuously from production traffic died on distribution shift and feedback contamination. Explainability never emerged from better algorithms, we are still bolting it on. Data scientists did not replace domain experts, the winning teams embedded modelers inside the business did. End-to-end automation gave way to decision support with a human in the loop wherever accountability mattered.

The pattern for evaluating today’s claims is the same. Discount anything promising to remove humans from workflows where liability and trust pull them back in. Discount continuous learning pitches that do not address regression detection. Discount framings where the model is the product rather than the system around it. Take seriously the unglamorous work on evaluation, observability, and feedback loops, because that is where durable value compounds

If you’re a C-suite executive at an enterprise trying to figure out where to invest in AI, what questions should you be asking your product and engineering teams to separate genuine opportunities from hype?

The questions that separate signal from hype are unglamorous, which is why exec conversations usually skip them. What is the cost per decision and the value per decision, and does the math survive at production volume. What is the baseline you are actually beating, a well-tuned rules engine or a simpler model rather than doing nothing. How do you know when the system is wrong and what happens when it is, because if the answer involves the word eventually the system is not production ready. Where does the human stay in the loop, and is that by design or by accident, because workflows that quietly require human cleanup are not automated, they are subsidized by labor you have not accounted for. What does your evaluation infrastructure look like and how often does it run against real production traffic.

The meta-question underneath all of these is whether your team is building a system or shipping a model. Serious teams talk about evaluation, observability, fallback paths, and unit economics.

Based on your experience building AI systems that process billions in commerce and serve millions of users, what’s the most common mistake organizations make when they’re trying to scale AI from pilot to production?

The most common mistake is treating the pilot as a smaller version of production, when it is actually a fundamentally different problem. Pilots succeed on curated data, motivated users, forgiving latency, and a team that personally babysits every edge case. Production fails on none of those being true. The model that hit 92 percent accuracy in the pilot now sees traffic from segments that were not in the training set, latency budgets that do not tolerate the prompt chain you built, users who will not tell you when the answer is wrong, and a long tail of inputs that your evaluation harness never imagined. Organizations get blindsided by this because the pilot metrics look great right up until launch, and then quality degrades in ways nobody instrumented for.

The second order mistake, which compounds the first, is underinvesting in the operational layer because it does not demo well. Evaluation infrastructure, observability, fallback paths, drift detection, human review queues, and the unsexy plumbing that catches problems before customers do. None of it shows up in a board deck, all of it determines whether you still have a product in six months. The teams that scale successfully treat the model as maybe 20 percent of the work and the system around it as the other 80, and they start building that system during the pilot rather than after. The pattern is so consistent that I now treat the ratio of model work to systems work in a team’s roadmap as the single best leading indicator of whether they will make it to production at scale.

  • Pallavi Singal is the Vice President of Content at ztudium, where she leads innovative content strategies and oversees the development of high-impact editorial initiatives. With a strong background in digital media and a passion for storytelling, Pallavi plays a pivotal role in scaling the content operations for ztudium's platforms, including Businessabc, Citiesabc, and IntelligentHQ, Wisdomia.ai, MStores, and many others. Her expertise spans content creation, SEO, and digital marketing, driving engagement and growth across multiple channels. Pallavi's work is characterised by a keen insight into emerging trends in business, technologies like AI, blockchain, metaverse and others, and society, making her a trusted voice in the industry.

Follow us on Google

Choose IntelligentHQ as one of your Preferred Sources to see more of our latest stories in Google.

Fill out the form below to request your copy.

Name(Required)