HomeFeatured build

Truefit.ai

A real-time AI voice interview platform. Candidates talk to an AI interviewer live — full duplex audio, sub-second turn-taking, tool-calling agent logic — and get structured, data-driven feedback instead of a generic score.

PythonFastAPIWebRTC (aiortc)Gemini Live APIPostgreSQLRedisNext.jsTypeScript
Architecture

Two channels, one session

Voice interviews need audio that survives real network jitter — a naive WebSocket audio pipe means every dropped or late packet blocks everything behind it in the same stream (head-of-line blocking), which is exactly the failure mode you can't have mid-interview.

Truefit splits the session in two: a WebSocket channel carries session signaling — state transitions, tool calls, control messages — while all audio moves over WebRTC/SRTP. SRTP is UDP-based and designed to tolerate loss and reorder without stalling the stream, so a bad packet degrades quality for an instant instead of freezing the whole call. That split was a deliberate trade — more moving parts to run and debug, in exchange for audio that stays responsive under real conditions instead of degrading gracefully in a lab demo and badly in production.

Production incident

The bug that blocked launch

Before launch, interview sessions would occasionally lock into an infinite response loop — the agent kept responding to itself. Two independent bugs were compounding:

  • A greedy queue drain. The audio consumer was pulling everything available off the queue in one pass, collapsing up to 40 discrete audio chunks into a single call to the model.
  • A clock reset. Frame pacing was derived from a timestamp that got reset mid-session, so the pacing logic computed nonsensical intervals for those merged chunks.

Together, the merged chunk landed at the model with corrupted timing, produced a response that re-entered the same consumer path, and the loop fed itself. Neither bug alone reproduced it reliably, which is what made it expensive to isolate — it only showed up when queue backlog and a clock boundary lined up at the same time. The fix was to bound the drain to a fixed chunk budget per cycle and re-derive pacing from a monotonic clock instead of one that could be reset mid-flight — closing both failure paths rather than papering over the symptom.

Agent design

A state machine, not a prompt

The interview agent is structured as a tool-calling, domain-driven state machine sitting behind abstract ports — the agent logic doesn't know or care whether it's talking to Gemini Live or OpenAI's realtime API underneath. Swapping providers is a port implementation, not a rewrite of the interview logic, scoring, or conversation state.

A custom AudioBridge handles PCM resampling and asyncio-based concurrency, so a single session can manage multiple audio participants without the state machine having to reason about the transport layer at all.

This is the deepest technical work I've shipped. Happy to walk through the architecture, the incident, or the code in a call.

Schedule a call