Friday, October 2, 2026

How Pudding works under the hood

The Pudding Team

Why a hiring product should explain its plumbing

When a candidate opens a Pudding session, shares their screen and starts talking, a lot has to go right in the first two seconds. Audio has to arrive without dropping. The agent has to hear the end of a sentence and not the middle of one. It has to see the screen well enough to ask about the thing the candidate just clicked, and it has to do all of that for a finalist you are paying an honorarium to, who will judge your company by how the session felt.

Most hiring tools never tell you what they run on, and most of the time you shouldn't care. Realtime is different. A recorded Loom can buffer. A form can be slow. A live, two-way work demo either feels like a conversation or it doesn't, and the difference is almost entirely infrastructure.

So here is what Pudding is built on, why we chose it, and how it was built. None of this is a secret. Anyone with a browser's developer tools open can see most of it. The reason to write it down is that hirers deserve to know what they're trusting, and candidates deserve to know what's listening.

The problem: a work demo is a conversation, not a form

The standard way to put AI in a product is request and response. The user types, the model thinks, the model answers. That shape works for a chatbot and it fails completely for an interview.

A Pudding session has to run the other way around. The candidate talks continuously. The screen changes continuously. The agent has to listen to both streams at once, decide when a thought has finished, ask a question that refers to what's on screen right now, and then shut up while the candidate answers. If any step waits for the previous one to finish, the pauses add up and the candidate starts talking to a wall.

Two things make that hard. The first is transport. Browser audio and screen video have to reach a server reliably over whatever Wi-Fi the candidate has, with latency low enough that an interruption lands mid-sentence instead of three seconds later. The second is the model. It has to take audio in and put audio out without a transcription step in the middle, because every hop between speech and text and back adds delay and strips out tone.

We didn't build either of those from scratch. A solo-founded company that tried to would still be building. Instead we picked the two pieces of infrastructure that the largest voice products in the world already run on, and spent our time on the part that is actually ours: what the agent asks, when it asks it, and how the result gets back to the hiring team.

Transport: LiveKit

The audio and screen video in a Pudding session travel over LiveKit, an open-source WebRTC stack with a managed cloud on top. If you've used voice mode in ChatGPT, you've used it. OpenAI built Advanced Voice on LiveKit Cloud, and the same framework now carries realtime audio for a long list of companies you've heard of and thousands you haven't.

What that buys a Pudding session, concretely:

  • The candidate's browser connects to the nearest edge, not to our server. LiveKit routes media over its own global network, so a candidate in Lagos and a candidate in Lisbon both get a connection that holds up. We don't operate data centers, and with LiveKit we don't need to.
  • The agent is a participant in the room, like a person. LiveKit's Agents framework treats our AI interviewer as another attendee that can subscribe to the candidate's microphone and screen, publish its own voice, and get told when someone starts or stops talking. Turn detection, interruption handling and voice activity detection are handled by the framework rather than by code we wrote at 2 a.m.
  • Screen share is a first-class stream. The candidate shares a window or a screen through the normal browser prompt. No plugin, no download, no screen-recording software with a dubious installer.
  • It is open source under Apache 2.0. If LiveKit the company disappeared tomorrow, the server and the framework would still exist and we could self-host them. That matters more for a small vendor than a large one. Hirers are right to ask what happens to a startup's product if its suppliers change, and the answer here is: nothing dramatic.

The honest trade-off is that we don't control the media path end to end. We traded that control for a transport layer that has already handled billions of calls and that we could never have matched with our own engineering. That is the correct trade for a hiring product, where the session has to simply work and nobody is impressed by a hand-rolled WebRTC stack.

The interviewer: Gemini, in realtime

The voice asking questions in a Pudding session is Google's Gemini, running through the Gemini Live API. This is not the Gemini you type to. It is the realtime, native-audio version: audio goes in, audio comes out, and video frames from the candidate's screen are streamed in alongside the audio so the model is looking at the same thing the candidate is describing.

Why native audio matters for an interview, in plain terms. The older way to build a voice agent was a chain: transcribe speech to text, send the text to a model, turn the reply back into speech. Three vendors, three round trips, and a reply that arrives a beat too late with the flat delivery of a text-to-speech engine. A native-audio model collapses that chain. It hears the hesitation before a candidate says "I'm not sure that was the right call," and it can ask about it. It stops talking when the candidate interrupts. The exchange on our home page, where the agent says "I see you chose that framework. Now, how did you decide to use this framework for this audience?", is a transcript of what that looks like in practice, not a script we wrote.

The second reason is the screen. Gemini's realtime models take live video as an input, which is what lets the agent ask about the Jira board or the Figma frame the candidate just opened rather than a generic list of competency questions. A work demo without that is just a phone screen with extra steps.

The third reason is the one hirers tend to ask about: adoption. Gemini's realtime stack is no longer an experiment. It is generally available on Google Cloud's enterprise platform, the native-audio models have been through several public releases this year, and LiveKit ships a first-party Gemini plugin, which means the two halves of our stack were designed to talk to each other. We didn't glue them together. We configured a connection that Google and LiveKit already maintain.

What the model does not do is decide anything. It listens, inquires and records. Scoring and the hiring decision stay with the humans on your team, and the transcript and recording are there so they can check the agent's work as easily as the candidate's.

How it was built, since you're going to ask

There is a fair question underneath every "what are you built on" conversation in 2026, and it's whether the product was built at all or merely prompted into existence. Hirers are allowed to ask. Pudding exists because we think the artifact alone proves less every month and the reasoning behind it is what counts, so it would be strange not to answer the same question about ourselves.

Pudding was written by a person, with AI copilots from Anthropic and Google in the loop. Claude and Gemini drafted code, caught bugs, and argued with us about architecture. They did not decide the architecture. Every design choice in this post, from using LiveKit rooms instead of raw WebRTC to keeping scoring out of the model's hands, was made by a human who understood the trade-off, and every line of code that implements it was read by that same human before it shipped. Not sampled, not spot-checked. Every line. The copilots were managed, in the ordinary sense that a senior engineer manages a fast junior one: give it a scoped task, review what comes back, reject what's wrong.

We say this partly because it's true and partly because it's the same standard we ask of candidates. A Pudding session doesn't penalize anyone for using Claude Code or Cursor to produce their work sample. It asks them to explain it, live, to something that will notice if they can't. We'd fail our own interview if the answer to "why did you build it this way" were a shrug.

It also has a practical consequence for you. When something breaks at 9 a.m. on the morning your finalist is scheduled, there is a person who knows why every piece is where it is, and that person is the one who answers the support email.

What this means if you're hiring

Three things, and then we'll stop talking about ourselves.

The session will hold up. The media path is the same one that carries ChatGPT's voice mode. The model is the same realtime stack Google sells to enterprises. The part we built sits on top of those and is small enough to be understood by the people who run it.

The model is replaceable, and we'll tell you when we replace it. We chose Gemini because its native-audio and live-video support was the best fit for a screenshare interview. That market moves fast, and if a better fit appears we'll switch. Nothing in how a hirer configures an engagement, reviews a submission or pays an honorarium depends on which model is behind the voice.

You can check all of this. Open the self-guided demo, share your screen, and talk to the agent. If the session feels like a conversation, the infrastructure did its job. If it doesn't, you'll know within thirty seconds, which is more than most vendor pages give you.

Work is the supersignal. The plumbing is just what carries it.

Sources