S
ACTIVEFull-stack EngineerProfessionalProduction

VoiceCon

Android-first AI voice calling: a caller's live speech is converted in real time into an authorized voice profile on a GPU and delivered to the recipient over a normal phone call.

KotlinJetpack ComposeFastifyPrismaPostgreSQLTwilioPythonPyTorchSeed-VC

[ 01 ] · Context

Problem

  • Changing a voice during a phone call has to happen in real time: the converted audio must reach the recipient with conversational latency, over a normal phone line, from a phone app.
  • Twilio's media streams return audio only to the call leg they run on, so a single leg that dials the phone network cannot carry converted audio to the recipient.
  • Voice conversion is gated: a caller may only speak as a voice profile they are authorized to use, and every conversion minute has to be metered against an active package.

[ 02 ] · Response

Solution

VoiceCon splits the product into a control plane and a media plane. A Fastify + Prisma backend on PostgreSQL owns accounts, voice profiles, packages, credits, and call authorization; a Python gateway handles realtime audio over Twilio Media Streams; and a GPU worker runs Seed-VC inference. The backend never touches audio packets, and the gateway and worker never decide billing or authorization: they ask the backend over HMAC-signed internal APIs.

Each call is a two-leg bridge. The Android app places a VoIP leg into Twilio, the gateway streams that audio to the GPU worker as PCM, and the converted audio is sent out on a second leg that dials the recipient on the phone network. The recipient's voice returns to the caller as passthrough.

Conversion fails closed: if the worker errors, both legs end instead of forwarding the caller's raw voice.

[ 03 ] · My work

My contribution

  • I build across the stack: the Kotlin and Jetpack Compose Android app (voice lab, packages, profile and KYC flows), the backend API and its tests, the React admin panel, and the deployment and migration tooling.
  • I wrote the access rules around conversion (an active package plus callable minutes) and the user-facing messaging that explains why conversion is unavailable, with a shortcut to the packages screen.

[ 04 ] · Structure

Architecture

Android ──VoIP──▶ Twilio leg A ──Connect/Stream──▶ Gateway ──PCM 16k──▶ GPU worker (Seed-VC)
                                                       │◀──converted PCM──┘
Recipient ◀─PSTN── Twilio leg B ◀──Connect/Stream──────┘   (A converted → B)
                   leg B audio ──▶ Gateway ──▶ leg A   (recipient voice → caller, passthrough)

Backend (Fastify + Prisma + PostgreSQL): accounts, voice profiles, billing ledger, call authorization
Admin panel (React + Vite)  ·  HMAC-signed internal APIs between gateway, worker, and backend

[ 05 ] · Trade-offs

Key engineering decisions

  1. 01

    Control plane and media plane stay separate

    The backend decides who may convert and how many minutes remain; the gateway and worker only move and transform audio. Keeping billing and authorization out of the audio path means a slow database call can never glitch a live call.

  2. 02

    Fail closed on conversion errors

    If voice conversion fails mid-call, both legs end with an explicit error. The caller's unconverted audio is never forwarded to the recipient.

  3. 03

    PostgreSQL is the only source of truth

    Credits and transactions live in the database, and call deadlines are derived from stored timestamps so a reconciler can enforce them even if an in-process timer is lost. Redis was deliberately left out until a second instance needs shared rate limits.

  4. 04

    Pick the GPU from a benchmark

    The RunPod GPU was chosen from a recorded benchmark of candidate cards, and the worker loads the model once at startup so a call never pays model-load latency.

[ 06 ] · Under the hood

Engineering highlights

The two-leg bridge
Android ──VoIP──▶ Twilio leg A ──<Connect><Stream>──▶ Gateway ──PCM16k──▶ GPU worker (Seed-VC)
                                                          │◀──converted PCM──┘
Recipient ◀─PSTN── Twilio leg B ◀──<Connect><Stream>──────┘   (A converted → B)
A media stream returns audio only to its own leg, so converted audio is carried to the recipient on a second leg that the backend dials only after leg A is authorized and a GPU session is ready.

[ 07 ] · Outcomes

Results

6
layers: app, API, admin, gateway, GPU worker, database
2
Twilio legs bridged per call
0
raw caller audio forwarded on failure
HMAC
signed internal service APIs

[ 08 ] · Where it's going

Status & next steps

  • A shared rate-limit store and a Postgres advisory lock for the reconciler, to run more than one backend instance.
  • Shared worker and session state so several gateways can schedule across the same GPU pool.

[ 09 ] · Keep exploring

More projects

InvoicePilot: AI-powered invoice processing platform that extracts, validates, and analyzes invoices while detecting duplicates and anomalies to streamline financial document workflows.
ProfessionalProduction
Featured

AI-powered invoice processing platform that extracts, validates, and analyzes invoices while detecting duplicates and anomalies to streamline financial document workflows.

FastAPIPythonMindee OCRPostgreSQLRedis StreamsCloudflare R2
Back to all projects