McDonald's drive thru AI upgrade — will Archy work?

Archy is back in a limited US pilot, but McDonald's has disclosed no accuracy, intervention, cost, or rollout…

McDonald's drive thru AI upgrade — will Archy work?

The short version

  • McDonald's drive-thru AI upgrade is a limited US Archy test with no public rollout timetable.
  • The previous IBM pilot ended after roughly three years, with software switched off at every participating restaurant.
  • Archy must preserve orders during human handoffs and prove accuracy and net transaction costs before broad deployment.

McDonald’s is testing another drive-thru AI without publishing its accuracy, rollout date, or even the restaurants involved. Naturally, the internet has already promoted it to robot cashier-in-chief. The McDonald’s drive-thru AI upgrade is a limited US test of ArchIQ and its “Archy” voice assistant. The company unveiled the system in June but never explained how Archy processes speech, handles an ambiguous order, or summons an employee when the conversation goes sideways.

Those details decide whether the thing works. Ordering a customized burger should not become an oral exam administered through a speaker that survived three Midwestern winters, especially while three cars are honking and somebody’s diesel pickup is cosplaying as industrial machinery. I once spent several apologetic minutes outside Barstow removing onions, restoring pickles, and changing a drink while the order screen preserved every decision like evidence for a future trial.

The employee fixed it in seconds. Any useful drive-thru AI has to fail that gracefully.

IBM’s order taker met actual customers

The IBM pilot began as a test of whether voice-activated software could speed service and simplify restaurant operations. More than a hundred US locations eventually participated, giving the Automated Order Taker far more exposure than the usual conference demo with perfect Wi-Fi and a product manager speaking like airport security. Actual drive-thrus supplied regional accents, menu substitutions, and customers changing their minds halfway through a sentence. They also supplied toddlers screaming from the back seat. Enterprise benchmarks rarely include that feature.

Bar chart comparing current figures against their baselines: dual-3090 DFlash2 greedy decode… 474 tok/s versus 432 tok/s, single-stream prose decode on the… 85 tok/s versus 98 tok/s, single-user speculative decode at long… 32 tok/s versus 68 tok/s, steady-state aggregate decode throughput… 1100 tok/s versus 46 tok/s.

An automated order taker has a brutally unforgiving job. Audio enters through a battered outdoor microphone, then speech-recognition software must turn it into language the ordering system can interpret. That interpretation has to match products available at the restaurant, apply modifiers, and notice when I have changed my mind for the fourth time because commitment scares me. If two interpretations remain plausible, the system has to ask a useful follow-up question rather than confidently invent lunch. The assembled basket must then reach the transaction system without losing quantities or substitutions. Only after all of that does the kitchen receive a ticket. Perfect transcription still fails if “one large fries” becomes two orders farther downstream.

IBM and McDonald’s said the test produced substantial benefits for customers and restaurant crews. Their statement included no accuracy result, service-time comparison, or cost baseline against a human order taker. “Substantial” was doing cardio there.

The pitch made sense: software handles routine exchanges while an employee prepares food or helps elsewhere. But the demo is merely the opening act. Somebody still has to maintain menu integrations, investigate failures, and update the system whenever a new product launches. Franchisees need support when the bot starts freelancing during lunch. Employees still have to rescue customers trapped in a loop about dipping sauce.

After roughly three years, the review ended with IBM’s technology being switched off at every participating restaurant by the July 2024 deadline. No expansion followed.

Mason Smoot, McDonald’s US chief restaurant officer, left the category open in his message to franchisees:

While there have been successes to date, we feel there is an opportunity to explore voice ordering solutions more broadly.

My translation: voice ordering still had appeal. This implementation had used up its lives.

Archy remains a very private experiment

At its Las Vegas convention in June 2026, the company presented ArchIQ and Archy—two years after ordering the IBM pilot shut down. Archy is the customer-facing voice assistant, while ArchIQ is the restaurant platform around it. CIO described the technology as Google-backed, although the exact relationship remains undisclosed. I cannot tell whether Archy uses Google Cloud, another vendor, technology carried over from elsewhere, or some combination of the three.

Karine Brunet described the reality of interconnected technology ecosystems:

Today’s organisations operate in highly interconnected technology ecosystems where complete independence is rarely achievable.

So far, all we know about the deployment is that it involves a limited number of American restaurants. We do not know how many or which ones. There is no rollout timetable and no published threshold that would trigger an expansion.

A serious pilot could still teach the company plenty, but only if it captures what happens after a customer starts speaking. The test would need current local prices and product availability, not a generic national menu. It would need to track whether spoken phrases become the correct items and modifiers, and how often ambiguous requests require another question. Employee interventions matter too: can a worker inherit the existing basket, or does the customer have to start again? Completed transactions could then be compared with human-taken orders for delays, remakes, and other failures. Franchisees could weigh any labor saved against the software bill and the time crews spend rescuing conversations. That is how I would evaluate Archy; the company has disclosed none of those test procedures.

McDonald’s drive-thru speaker, cars, crew member, and service window beneath iconic wordmark in bright daylight.

The missing architecture matters because the friendly voice from the speaker is only one endpoint inside an old, noisy restaurant. Archy somehow needs current menu information and must know what that location can sell right now. When the ice cream machine enters its traditional state of spiritual unavailability, the assistant has to stop offering McFlurrys. Accepted requests need to become valid transactions with the correct quantities and customizations. Sold-out items and conflicting modifiers need some form of resolution. Uncertain orders need some route to an employee, ideally without throwing away the conversation and assembled basket. Otherwise, the lane backs up while everyone discovers new uses for the car horn. Available sources explain none of this, so anyone claiming to know Archy’s mechanism is guessing.

Secrecy during a pilot is normal. Nobody owes me a technical diagram whenever a microphone gets tested in Ohio. Franchisees do deserve credible operating numbers before anyone asks them to fund a rollout, though, and public talk of an “upgrade” deserves more evidence than convention slides.

The Google relationship requires the same restraint. A broad Google Cloud partnership covering generative AI, cloud services, and edge tools was announced in 2024. Two years later, reporting still does not explain precisely what Google built for ArchIQ. “Google-backed” is supported. “Google built Archy” is fan fiction.

Martin Merz framed digital resilience in terms of sovereignty:

La résilience numérique de l’Europe repose sur une souveraineté à la fois sécurisée et scalable.

The voice assistant will generate the TikTok clips. Whether anyone gets lunch depends on the boring integration work.

Accuracy gets expensive in the kitchen

CNBC cited two sources familiar with IBM’s technology who said accents and dialects created interpretation problems that hurt order accuracy. The fast-food chain declined to discuss specific technical issues. IBM described the system as fast and accurate under demanding conditions, while viral videos showed spectacularly weird orders. I trust TikTok with memes and recipes involving irresponsible amounts of burrata. Auditing restaurant software is beyond its remit.

BTIG analyst Peter Saleh told CNBC that his industry checks placed IBM’s accuracy in the low-to-mid 80% range, versus the minimum 95% he considered commercially viable. Those were analyst channel checks, not audited measurements released by either partner. Saleh also reported franchisee frustration and high operating costs. I would not tattoo his estimate onto an earnings model, but I would not ignore it either.

That gap explains the business problem. The system hears a request and produces an interpretation, which must be reconciled with the menu and everything the customer has already said. Clear requests can continue through the order flow. Poor audio, contradictory changes, or an unfamiliar dialect create uncertainty that needs to be contained. An employee taking over would need the existing basket plus the unresolved question because a blank screen wastes time and blocks the lane. If the uncertain order reaches the kitchen first, software uncertainty becomes physical food that must be discarded or remade. Nobody outside the pilot knows Archy’s control loop. Any safe design needs an equivalent way to stop doubt before somebody cooks it.

The economics follow the same path. Count labor saved and any additional orders processed, then deduct the software bill and employee intervention time. Mistakes consume ingredients and queue capacity while the crew untangles them. A voice assistant can look magical in a laboratory and still lose money one incorrect Quarter Pounder at a time.

An average accuracy score can also hide who absorbs the pain. Respectable overall performance could coexist with miserable results for customers from Naples, Lagos, or rural Louisiana. The real test is performance across accents, dialects, menu changes, and noisy drive-thru conditions. No such results have been published for Archy.

There is a fair argument for continuing: the IBM pilot apparently showed enough promise that voice ordering survived the breakup. I buy that. Ending one vendor partnership does not kill the category. Still, shutting down the software at every test restaurant raises the evidence bar for whatever comes next.

I want to see the orders Archy refuses to finish. Calling a human before doubt becomes a bag of unwanted Filet-O-Fish is much harder—and more valuable—than delivering confident nonsense.

AI market headlines tell us nothing about Archy

There is no disclosed connection between this pilot and an Anthropic IPO. A future listing could reveal plenty about Anthropic’s finances while telling us absolutely nothing about intervention rates inside a drive-thru. Available reporting names IBM in the previous pilot and describes Google-backed technology around ArchIQ. Anthropic never appears in that chain.

The same separation applies to agentic AI coding tools. Coding agents plan tasks, edit files, and retry failed tests inside a controlled workspace. Archy has to interpret live speech against a changing restaurant menu before the cars behind me begin a small revolution. Both products may use large models somewhere in their stacks. That shared ingredient tells us very little about how either behaves, much as a pizza oven and a blast furnace both get hot while only one belongs near my dinner.

Neither the MSFT earnings date nor the ORCL earnings date provides rollout guidance. Microsoft and Oracle earnings can reveal broad cloud demand. They cannot identify Archy’s suppliers or tell a franchisee what each completed order costs.

Here is the evidence chain I care about. Start with the named vendor and the restaurants included in the test. Then measure performance against the human process being replaced. Track interventions, failed orders, remakes, and time saved. Add the software and support costs that franchisees would actually pay. Archy currently gives us a product name, glossy footage, and confirmation of a limited American pilot. Restaurant identities, accuracy, intervention frequency, operating cost, and expansion criteria remain secret. Without those details, every grand rollout claim is guesswork wearing a headset.

My dated bet: Archy will not begin a broad US rollout before the end of 2027 unless franchisees first receive credible completed-order results against human staff and a real net cost per transaction.

Until then, the smartest voice in the drive-thru may still be the employee saying, “I’ve got it from here.”

Frequently asked questions

What is McDonald's drive thru AI upgrade?

McDonald's drive thru AI upgrade is a limited US test of ArchIQ and its customer-facing Archy voice assistant. McDonald's has not disclosed the participating restaurants, rollout timetable, accuracy, intervention frequency, operating cost, expansion criteria, or the precise role Google plays in the system.

Why did McDonald's end its IBM drive-thru AI pilot?

McDonald's ended the IBM Automated Order Taker pilot after roughly three years and switched off the technology at every participating restaurant by July 2024. The partners cited substantial benefits but published no accuracy, service-time, or cost baseline. Analyst channel checks also reported accuracy problems, franchisee frustration, and high operating costs.

Why do human handoffs matter for drive-thru AI?

A useful drive-thru AI handoff gives an employee the existing basket and unresolved question without forcing the customer to restart. Preserving quantities, substitutions, and context prevents uncertainty from reaching the kitchen, where an incorrect interpretation consumes ingredients, causes remakes, and slows the lane.

Sources

Related reading

Luca

Luca

Luca by the way is the personal blog of Los Angeles based entrepreneur Luca Capula. A true Italian who lives between Torino and LA.

More posts →