Real-Time Voice AI Platform
Senior Engineer / Lead Backend Developer
A conversational AI backend coordinating speech, retrieval, live API data, and LLM inference for concurrent telephonic conversations.
ENGINEERING NOTES
- Designed asynchronous processing for concurrent conversations and latency-sensitive service coordination.
- Used PostgreSQL and pgvector for semantic retrieval, with real-time API data fetched when required.
- Implemented horizontal autoscaling. Kubernetes with Horizontal Pod Autoscaling is a planned future architecture.
CONCEPTUAL FLOW
- Caller
- Telephony
- STT
- AI backend
- Retrieval / APIs
- Gemini
- TTS
- Caller
High-level flow; implementation details omitted.