- Contentassemblyai.com
DER vs. cpWER: why the standard diarization metric ranks systems backwards
AssemblyAI’s research demonstrates that the standard diarization metric DER (Diarization Error Rate) ranks speech-to-text outputs backwards compared to human judgment. Using real audio clips and produ…
- Featureassemblyai.com
AssemblyAI Universal-3.5 Pro benchmarks
AssemblyAI released comparative benchmarks for its Universal-3.5 Pro and Universal-3.5 Pro Realtime models, demonstrating lower word error rates (WER), missed entity rates, and diarization errors than…

- Integrationassemblyai.com
AssemblyAI integrates Universal 3.5 Pro Realtime with LiveKit for voice agents
AssemblyAI launched Universal 3.5 Pro Realtime, a next-generation streaming speech-to-text model, and integrated it into LiveKit’s voice agent framework via the `livekit-agents` SDK (v1.6+). The integ…
- Featureassemblyai.com
AssemblyAI Universal-3.5 Pro adds contextual and keyterms prompting for transcription
AssemblyAI introduced two new prompting methods for its Universal-3.5 Pro model to improve transcription accuracy in challenging audio scenarios. Contextual prompting allows users to provide natural-l…
- Featureassemblyai.com
AssemblyAI adds transcript summarization in open beta with chapter timestamps
AssemblyAI introduced a summarization feature for audio transcripts that splits output into timestamped chapters with headlines, available in open beta. Users can choose between bullet or paragraph fo…
- Launchassemblyai.com
Introducing the Sync API: transcripts in a single call
AssemblyAI introduced the Sync API, enabling one HTTP request to transcribe short audio clips with Universal-3.5 Pro in ~134 ms, eliminating polling, WebSockets, or chunking. The API targets use cases…

- Featureassemblyai.com
AssemblyAI adds real-time webhooks for voice agent lifecycle events
AssemblyAI introduced a webhook system for its voice agent API, enabling real-time notifications for session and call lifecycle events. Developers can subscribe to events like `session.started`, `call…
- Featureassemblyai.com
AssemblyAI Voice Agent API improves semantic turn detection and barge-in
AssemblyAI’s Voice Agent API now uses semantic end-of-turn detection and adaptive pacing by default, eliminating the need for manual tuning. The system waits for complete tool values (e.g., phone numb…
- Featureassemblyai.com
AssemblyAI Voice Agent API adds key terms for transcription accuracy
AssemblyAI introduced a new `input.keyterms` feature in its Voice Agent API to boost transcription accuracy for rare or domain-specific terms like brand names, proper nouns, and jargon. Developers can…
- assemblyai.com
AssemblyAI Voice Agent API adds parameter hints and execution modes for tools
AssemblyAI updated its Voice Agent API tools documentation to introduce parameter hints (enum, examples, pattern, format) that improve tool-calling accuracy and turn detection. Execution modes and pro…
- Integrationassemblyai.com
AssemblyAI enables custom OpenAI-compatible LLM endpoints for voice agents
AssemblyAI now allows voice agents to use custom OpenAI-compatible LLM endpoints instead of their managed model. Users can configure `base_url`, `model`, and `api_key` in the agent settings to route c…
- Featureassemblyai.com
AssemblyAI enables mid-stream configuration updates for streaming transcription
AssemblyAI introduced the ability to update streaming session parameters mid-session using an `UpdateConfiguration` message without reconnecting. Users can dynamically adjust accuracy/latency modes, t…
- Featureassemblyai.com
AssemblyAI launches Universal-3-5-Pro Streaming with new language steering and domain models
AssemblyAI updated its Streaming WebSocket API to introduce Universal-3-5-Pro Streaming as the default model, replacing the prior default. The update adds language steering via the `language_codes` pa…
- Featureassemblyai.com
AssemblyAI adds action items extraction in speech understanding beta
AssemblyAI introduced a new action items feature in beta that generates timestamped action items, quotes, and effort levels from audio transcripts. The feature supports US and EU regions, offers low (…
- Launchassemblyai.com
Universal-3.5 Pro: native code switching, our most accurate speaker diarization yet, and expanded language support
AssemblyAI released Universal-3.5 Pro, a flagship async speech-to-text model featuring native code-switching across 18 languages, the most accurate speaker diarization to date, and contextual promptin…

- Featureassemblyai.com
Contextual Awareness in Universal-3.5 Pro Realtime
AssemblyAI introduced contextual awareness in Universal-3.5 Pro Realtime, enabling the model to use conversation history, speaker context, and dynamic prompts to improve transcription accuracy in nois…
Track AssemblyAI on autopilot
- · Weekly AI brief: narrative summary of what shipped, every Monday 9 AM
- · Email or Slack alerts, or chat with the archive in your dashboard
- · Add AssemblyAI + up to 2 more competitors free, no credit card