
Cosellr
Rapid MvP / AI Voice Agents / Collaborative Agents / Platform Engineering / GCP / AWS / B2B SaaS
Executive brief
Cosellr needed voice agents that feel human in conversation while still behaving like reliable software—fast responses, safe tool usage, and consistent outcomes. We implemented a Pipecat-based real-time pipeline on top of Daily.co for voice transport, enabling low-latency turn-taking and multi-step workflows (qualification, scheduling, follow-ups, and CRM updates). The system includes guardrails for compliance and tone, plus tracing and evaluation foundations so conversational quality improves over time instead of drifting.
Median voice response latency (end-to-end)
0.9–1.8s
steady-state target
P95 voice response latency (non-tool turns)
2.5–4.0s
steady-state target
Call-to-meeting conversion lift
15–35%
within 6–10 weeks
Built low-latency, real-time voice agents using Pipecat + Daily.co optimized for natural turn-taking and tool-driven workflows.
Implemented end-to-end call workflows: qualification, objection handling, scheduling, follow-ups, and CRM note/field updates.
Introduced guardrails for policy, tone, and PII handling with measurable reductions in failure modes (hangups, dead-ends, hallucinated claims).
Established observability and eval foundations: call transcripts, event timelines, tool traces, and regression tests for conversational quality.
Engagement shape
Duration
12 wks
Phases
5
Deliverables
8
Median voice response latency (end-to-end)
0.9–1.8s
How it ran
Foundation-first build with rapid iteration loops and staged rollout across personas/campaigns.
Cosellr’s voice agents use Pipecat and Daily.co to run real-time sales calls with tool-backed reasoning, guardrails, and production-grade observability.
Starting position
Traditional sales automation stops at email and chat. Cosellr wanted voice agents that can hold real conversations, execute tool-backed actions (like scheduling and CRM updates), and maintain consistent behavior across campaigns. Early prototypes often struggled with latency, awkward turn-taking, inconsistent reasoning, and missing operational visibility—making it hard to scale responsibly.
Constraints we worked inside
5 in playEvery one of these ruled an easier option out. They are the reason the approach looks the way it does.
Low-latency requirements for natural turn-taking (avoid long silences and interruptions).
Tool correctness and safety (no hallucinated facts, controlled side effects).
PII handling constraints and guardrails suitable for real customer/prospect conversations.
Campaign/persona scalability (multiple scripts, tones, and workflows without bespoke engineering each time).
Operational observability: trace what happened on a call, why it happened, and what tools were invoked.
What had to be true to finish
5 goalsAgreed up front, so the engagement could be called finished on evidence instead of opinion.
Build real-time voice agents using Pipecat + Daily.co with production-grade reliability.
Enable end-to-end workflows: qualify → handle objections → schedule → follow up → update CRM.
Add guardrails for policy, tone, and PII handling with auditable behavior.
Implement observability and evaluation foundations to drive iterative improvement.
Create a scalable architecture to support multiple personas and campaigns safely.
Industry context
Rapid MvP
AI Voice Agents
Collaborative Agents
Platform Engineering
GCP
Outbound sales and customer engagement increasingly demand real-time, high-quality conversations that can reliably connect to business systems. Voice AI must balance natural interaction with strict governance, tool correctness, and measurable performance—otherwise it becomes a liability rather than a growth lever.
Approach
We built the voice agent system as a real-time pipeline: Daily.co for audio transport, Pipecat for orchestration and streaming control, and a tool layer for business actions (CRM, calendar, enrichment, follow-up automation). We focused heavily on latency, turn-taking ergonomics, safe tool usage, and operational visibility—so the system could scale beyond a prototype into a dependable product capability.
Delivery track
Segment width reflects the number of workstreams inside each phase — where the engagement actually spent its effort.
Phase 1 — Real-time voice pipeline and latency optimization
4 workstreams
Implemented Daily.co as the real-time audio transport layer.
Built Pipecat pipelines for streaming ASR/TTS and conversational turn management.
Established targets for latency and conversational pacing (silence handling, interruption policy, barge-in behavior).
Added fallback behaviors for partial failures (ASR degradation, tool timeouts, network jitter).
Phase 2 — Tool-driven call workflows (sales actions, not just talk)
4 workstreams
Designed structured conversation flows for qualification, discovery, objections, and next-step commitment.
Integrated scheduling flows (calendar availability checks, booking, confirmations).
Integrated CRM writebacks (notes, outcomes, structured fields, follow-up tasks).
Implemented follow-up automations (email/SMS triggers, sequences) based on call outcomes.
Phase 3 — Guardrails, safety, and PII controls
4 workstreams
Implemented policy constraints for what the agent can claim, promise, or do.
Added tone controls per persona/campaign (confidence, friendliness, formality).
Introduced PII handling rules: redaction where needed, restricted logging, and safe storage conventions.
Implemented escalation / handoff patterns for edge cases and high-risk situations.
Phase 4 — Observability and evaluation foundations
4 workstreams
Captured call event timelines (audio start/stop, interruptions, tool calls, failures, retries).
Logged tool traces and structured outcomes (qualified/not qualified, next step, scheduled meeting, objections).
Built evaluation datasets and regression checks for conversational quality over time.
Established dashboards for success metrics and failure-mode monitoring.
Phase 5 — Production hardening and scale-out
4 workstreams
Standardized deployment patterns for multiple personas/campaigns.
Introduced concurrency controls and cost controls (tool budgets, retry budgets).
Created runbooks for incident response and behavior drift.
Shipped iteration loops so product teams can improve flows without destabilizing production.
What the team kept
Artefacts handed over at the end of the engagement — owned and operable by Cosellr without us.
Pipecat-based real-time voice agent pipeline integrated with Daily.co
Turn-taking design and latency tuning (silence, interruption, barge-in policies)
Tool layer integrations: scheduling + CRM writebacks + follow-up automations
Conversation flow library: qualification, objections, next-step commitment
Guardrails package: policy constraints, tone controls, PII handling conventions
Observability foundation: call timelines, tool traces, structured outcomes
Evaluation foundation: baseline call sets, regression checks, quality scorecards
Production readiness package: runbooks, dashboards, scale patterns
Results
Cosellr’s voice agents evolved from “talking demos” to production-grade systems capable of completing real workflows. The platform supports natural real-time conversation, tool-backed actions, and measurable outcomes—while remaining governable and observable. This created a scalable foundation for voice-driven outbound and customer engagement.
Outcome ledger
Every figure we measured on this engagement, including the ones that are ranges rather than headlines.
Median voice response latency (end-to-end)
Estimate for natural turn-taking; varies by ASR/TTS and tool usage. Replace with telemetry (median/P95).
0.9–1.8s
steady-state targetP95 voice response latency (non-tool turns)
Estimate excluding long tool calls; replace with real P95 once instrumented.
2.5–4.0s
steady-state targetCall-to-meeting conversion lift
Estimated improvement vs baseline outbound processes when using consistent qualification + scheduling flows.
15–35%
within 6–10 weeksAgent-driven scheduling success rate
Estimate for calls where scheduling is attempted and prospect is willing; depends heavily on ICP and script.
60–80%
per scheduling attemptReduction in “dead-end calls” (no next step captured)
Estimated reduction from structured next-step capture and fallback behaviors (handoff, recap, scheduling).
30–55%
after flow + guardrail tuningDecrease in hallucinated/unsupported claims
Estimate driven by policy constraints, tool-based grounding, and strict refusal behaviors.
70–90%
after policy + tool groundingWhat changed day to day
The part of the result that never shows up in a dashboard, but is the reason the numbers held.
Sales teams get a voice channel that behaves like a reliable workflow engine, not an unpredictable chatbot.
Prospects experience faster, smoother calls with clear next steps and fewer awkward pauses or interruptions.
RevOps gains structured outcomes: consistent qualification fields, notes, tasks, and follow-ups pushed into CRM automatically.
Product teams can iterate on scripts and personas without destabilizing the platform due to guardrails and evals.
The system becomes continuously improvable: call data + traces + evals drive measurable quality gains over time.
Your turn
Cosellr’s voice agents use Pipecat and Daily.co to run real-time sales calls with tool-backed reasoning, guardrails, and production-grade observability. If that shape looks familiar, the first conversation is a working session, not a pitch.
Real-Time Voice Agents
Agent Platforms & Orchestration
Conversation Intelligence
Tooling Integration (CRM + scheduling + enrichment)
Reliability Engineering (latency, retries, fallbacks)
Security & PII Controls
What Cosellr got
Median voice response latency (end-to-end)
0.9–1.8s
steady-state target
P95 voice response latency (non-tool turns)
2.5–4.0s
steady-state target
Call-to-meeting conversion lift
15–35%
within 6–10 weeks