Engineering expertise

Voice AI & LLM products

Voice apps, speech intelligence & LLM integration engineering

My AI work is application engineering, not model training or data science. I integrate speech, audio-analysis and LLM systems into products and evaluation pipelines, combining model outputs with deterministic application logic when one model should not control the entire decision. The work spans capture and transcription, diarization, speaker-role isolation, acoustic affect, semantic reasoning, structured outputs, audio events, noise and silence diagnostics, latency/cost trade-offs, model serving, and production APIs.

What this looks like in practice

  • Built JournPad voice workflows around recording, speech-to-text, preserved audio, playback, and AI-assisted organization of spoken journal entries.
  • Built the AutoAce voice-analysis pipeline using ElevenLabs Scribe v2 for transcripts, diarization, word timestamps and audio-event tags; customer-speaker inference then isolates customer-only speech segments.
  • Served emotion2vec-plus-base on Modal using FunASR, ModelScope, PyTorch, torchaudio and FFmpeg to score acoustic affect over customer speech segments rather than treating the full call as one speaker.
  • Combined Gemini raw-audio analysis and structured semantic reasoning with Scribe diagnostics and emotion2vec acoustic evidence, then fused the signals with named deterministic rules for tone, intensity, background noise, audio quality, speaker overlap and long-silence classification.
  • Used FastAPI, Pydantic structured schemas, httpx/provider clients and asynchronous provider orchestration to expose validated batch audio analysis while tracking latency, model usage and accuracy/cost trade-offs.
  • Built image-understanding and recommendation workflows in AI Stylist plus LLM-assisted research, outreach and operational automation around other products.

Core technologies

Voice appsSpeech-to-textWhisperElevenLabs Scribe v2Speaker diarizationWord timestampsGemini 3.6 FlashGoogle GenAI SDKStructured LLM outputsemotion2vecFunASRModelScopePyTorchtorchaudioFFmpegModalFastAPIPydanticAudio event detectionSpeaker-overlap detectionSilence detectionDeterministic signal fusionLLM reasoningOpen-source models

Project evidence

Products that demonstrate this capability

Each project below links to a deeper case study covering product scope, engineering ownership, supporting systems, and production evidence.

SaaS · Case study

JournPad: Voice Journal

Voice-first journaling app that preserves the original recording, transcribes spoken entries, and uses AI to generate useful titles, summaries, and categories. Entries can be searched and revisited by date, linked to goals, supported by reminders and prompts, and protected with account, deletion, and optional biometric controls.

Read case study

SaaS · Case study

AI Stylist

Mobile wardrobe-management and outfit-recommendation product. Users create wardrobes, upload clothing photos for AI-assisted category, tag, and description analysis, and generate one-time, daily, or weekly outfit suggestions that can use precise local weather; recommendation history is stored locally in SQLite and designed to sync with the cloud.

Read case study

Enterprise · Case study

RentPayor

Rent collection and reconciliation software for landlords and property managers. RentPayor creates rent invoices, lets tenants pay KES rent through an invoice-linked M-Pesa flow, automatically reconciles confirmed payments, and keeps partial balances, carried-forward credits, receipts, leases, units, tenants, and manual payment records in one rent ledger.

Read case study

Let's work together

Looking for an engineer who can own the product beyond a single layer?

I'm interested in senior product engineering work across web, mobile, backend, and integrations—especially teams shipping real products end-to-end.