Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
58 changes: 58 additions & 0 deletions technologies/assemblyai/assemblyai-guardrails.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
---
title: "Guardrails API"
author: "AssemblyAI"
authorUsername: "stevekimoi"
description: "AssemblyAI's Guardrails API adds safety and compliance controls to audio and transcripts, including PII text and audio redaction, profanity filtering, and content moderation."
---

# Guardrails API

The [Guardrails API](https://www.assemblyai.com/products/guardrails) adds safety, privacy, and compliance controls to AssemblyAI transcripts and audio. It covers PII redaction in both transcript text and the underlying audio, profanity filtering, and content moderation, applied as add-ons to a transcription request.

| General | |
|---------|--|
| **Developer** | AssemblyAI |
| **Type** | Safety and compliance API |
| **License** | Commercial API |
| **Documentation** | [assemblyai.com/products/guardrails](https://www.assemblyai.com/products/guardrails) |

---

## Core Features

- **PII text redaction**: removes personally identifiable information from transcript text.
- **PII audio redaction**: redacts personally identifiable information directly in the audio file.
- **Profanity filtering**: filters profane language out of transcript output.
- **Content moderation**: flags or filters unsafe or policy-violating content.

---

## Pricing

Guardrails features are billed as add-ons on top of the base Speech-to-Text rate:

| Add-on | Rate |
|--------|------|
| Profanity Filtering | +$0.01/hour |
| PII Text Redaction | +$0.08/hour |
| Content Moderation | +$0.15/hour |

Source: [assemblyai.com/pricing](https://www.assemblyai.com/pricing)

---

## Use Cases

- **Regulated industries**: healthcare and financial services use cases that require PII redaction before storing or sharing transcripts.
- **Customer support and contact centers**: content moderation and profanity filtering on recorded or live calls.
- **Compliance workflows**: redacting sensitive information from audio archives before retention or sharing.

---

## Ecosystem and Integrations

Guardrails add-ons apply to output from both [Speech-to-Text](/tech/assemblyai/assemblyai-speech-to-text) and [Universal-Streaming](/tech/assemblyai/assemblyai-streaming-speech-to-text), and can be combined with [Speech Understanding](/tech/assemblyai/assemblyai-speech-understanding) add-ons on the same request.

---

Read the [Guardrails docs](https://www.assemblyai.com/products/guardrails) to add safety and compliance controls to an existing transcription request.
59 changes: 59 additions & 0 deletions technologies/assemblyai/assemblyai-llm-gateway.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,59 @@
---
title: "LLM Gateway"
author: "AssemblyAI"
authorUsername: "stevekimoi"
description: "LLM Gateway is AssemblyAI's OpenAI-compatible API for calling frontier language models like Claude, GPT, and Gemini inside voice AI workflows, replacing the deprecated LeMUR product."
---

# LLM Gateway

[LLM Gateway](https://www.assemblyai.com/products/llm-gateway) is AssemblyAI's unified, OpenAI-compatible API for calling frontier language models inside voice AI workflows. It replaces LeMUR ("Leveraging Large Language Models to Understand Recognized speech"), which is deprecated and stops working March 31, 2026.

| General | |
|---------|--|
| **Developer** | AssemblyAI |
| **Type** | LLM gateway / unified API for voice AI workflows |
| **License** | Commercial API |
| **Documentation** | [assemblyai.com/products/llm-gateway](https://www.assemblyai.com/products/llm-gateway) |

---

## Core Features

- **OpenAI-compatible API**: a single interface for calling multiple LLM providers.
- **Multi-provider failover**: requests can fail over between providers.
- **Tool calling and JSON output**: supports structured output and tool/function calling.
- **Zero-data-retention option**: available for workflows with stricter data handling requirements.
- **Supported providers**: OpenAI, Anthropic, Google Gemini, Alibaba Qwen, and Moonshot Kimi.

---

## Pricing

LLM Gateway is billed per 1M tokens, with separate input and output rates:

| Model | Input (per 1M tokens) | Output (per 1M tokens) |
|-------|------------------------|--------------------------|
| Claude 4.5 Sonnet | $3.00 | $15.00 |
| GPT-5.1 | $1.25 | $10.00 |
| Gemini 3.5 Flash | $1.50 | $9.00 |

Source: [assemblyai.com/pricing](https://www.assemblyai.com/pricing)

---

## Use Cases

- **Transcript summaries and chapter markers**: generate summaries directly from transcript output.
- **Structured data extraction**: pull structured fields out of call transcripts.
- **Live LLM enhancement**: apply LLM processing to transcript streams as they arrive.

---

## Migrating from LeMUR

Existing LeMUR integrations need to move to LLM Gateway before LeMUR stops working on March 31, 2026. AssemblyAI publishes a [migration guide](https://www.assemblyai.com/docs/llm-gateway/migration-from-lemur) covering the API changes.

---

Read the [LLM Gateway docs](https://www.assemblyai.com/products/llm-gateway) or the [migration guide](https://www.assemblyai.com/docs/llm-gateway/migration-from-lemur) to get started.
65 changes: 65 additions & 0 deletions technologies/assemblyai/assemblyai-speech-to-text.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,65 @@
---
title: "Speech-to-Text"
author: "AssemblyAI"
authorUsername: "stevekimoi"
description: "AssemblyAI's Speech-to-Text API transcribes pre-recorded audio and video across 99 languages using the Universal-3.5 Pro model, with diarization, PII redaction, and native code-switching."
---

# Speech-to-Text

AssemblyAI's [Speech-to-Text API](https://www.assemblyai.com/docs/getting-started/universal-3-5-pro) transcribes pre-recorded audio and video across 99 languages. The current flagship model is Universal-3.5 Pro, released July 7, 2026 as the successor to Universal-3 Pro. AssemblyAI also keeps Universal-2 available as a lower-cost legacy tier.

| General | |
|---------|--|
| **Release date** | July 7, 2026 (Universal-3.5 Pro) |
| **Developer** | AssemblyAI |
| **Type** | Asynchronous speech-to-text API |
| **License** | Commercial API |
| **Documentation** | [assemblyai.com/docs](https://www.assemblyai.com/docs) |
| **GitHub** | [AssemblyAI/assemblyai-python-sdk](https://github.com/AssemblyAI/assemblyai-python-sdk) |

---

## Core Features

- **99 languages**: transcription coverage for pre-recorded audio and video.
- **Native code-switching**: audio that moves between languages mid-recording is transcribed without selecting a language in advance, across 18 languages: English, Spanish, French, German, Italian, Portuguese, Arabic, Danish, Dutch, Finnish, Hebrew, Hindi, Japanese, Mandarin, Norwegian, Swedish, Turkish, and Vietnamese.
- **Speaker diarization**: identifies and labels individual speakers in a transcript, available as an add-on.
- **PII redaction**: removes personally identifiable information from transcript text or the underlying audio.
- **Contextual prompting**: developers can pass domain context (e.g. medical terminology) with a request; in AssemblyAI's internal healthcare testing this reduced missed medical terms by 31%, and customer Metaview reported roughly 47% fewer low-confidence tokens.
- **Synchronous transcription**: a separate mode for short clips up to 120 seconds, returning a transcript in a single request/response instead of polling or webhooks.

---

## Benchmarks

On AssemblyAI's published code-switching benchmark (normalized WER), Universal-3.5 Pro scored 7.69%, compared to ElevenLabs Scribe v2 (8.77%), Deepgram Nova-3 (12.22%), and OpenAI GPT-4o (44.58%). Its predecessor, Universal-3 Pro, scored 9.07% on the same test.

On speaker diarization (cpWER), Universal-3.5 Pro averaged 30.17%, compared to Deepgram Nova-3 English (37.92%), ElevenLabs Scribe v2 (35.26%), and Gladia (36.87%).

Source: [Universal-3.5 Pro announcement](https://www.assemblyai.com/blog/universal-3-5-pro-async)

---

## Pricing

| Model | Rate | Notes |
|-------|------|-------|
| **Universal-3.5 Pro** | $0.21/hour | current flagship async model |
| **Universal-2** | $0.15/hour | legacy tier, still supported |

Speaker diarization adds $0.02/hour. New accounts get $50 in free credit at signup with no credit card required.

Source: [assemblyai.com/pricing](https://www.assemblyai.com/pricing)

---

## Tools and Resources

- **[Python SDK](https://github.com/AssemblyAI/assemblyai-python-sdk)**: official client for the Speech-to-Text API.
- **[Node.js/TypeScript SDK](https://github.com/AssemblyAI/assemblyai-node-sdk)**: official JavaScript and TypeScript client.
- **[API Reference](https://www.assemblyai.com/docs)**: full documentation for transcription, diarization, and redaction endpoints.

---

Start with the [$50 free credit](https://www.assemblyai.com/pricing) and the [Speech-to-Text docs](https://www.assemblyai.com/docs/getting-started/universal-3-5-pro) to send your first transcription request.
64 changes: 64 additions & 0 deletions technologies/assemblyai/assemblyai-speech-understanding.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
---
title: "Speech Understanding API"
author: "AssemblyAI"
authorUsername: "stevekimoi"
description: "AssemblyAI's Speech Understanding API turns transcripts into structured intelligence: sentiment analysis, entity detection, topic detection, summarization, translation, and speaker ID."
---

# Speech Understanding API

The [Speech Understanding API](https://www.assemblyai.com/products/speech-understanding) layers structured intelligence on top of AssemblyAI transcripts. It covers sentiment analysis, entity detection, topic detection, summarization, translation, speaker identification, custom formatting, key phrase extraction, and auto chapters, all applied as add-ons to a transcription request.

| General | |
|---------|--|
| **Developer** | AssemblyAI |
| **Type** | Audio/transcript intelligence API |
| **License** | Commercial API |
| **Documentation** | [assemblyai.com/products/speech-understanding](https://www.assemblyai.com/products/speech-understanding) |

---

## Core Features

- **Sentiment analysis**: scores sentiment across a transcript or per speaker turn.
- **Entity detection**: identifies named entities such as people, organizations, and locations.
- **Topic detection**: classifies transcript content by topic.
- **Summarization**: generates summaries and auto chapters from a transcript.
- **Translation**: translates transcript output into other languages.
- **Speaker identification**: attributes transcript segments to individual speakers.
- **Custom formatting and key phrase extraction**: formats output and surfaces key phrases for downstream use.

---

## Pricing

Speech Understanding features are billed as add-ons on top of the base Speech-to-Text rate:

| Add-on | Rate |
|--------|------|
| Translation | +$0.06/hour |
| Entity Detection | +$0.08/hour |
| Sentiment Analysis | +$0.02/hour |
| Topic Detection | +$0.15/hour |

Source: [assemblyai.com/pricing](https://www.assemblyai.com/pricing)

---

## Use Cases

- **Meeting intelligence**: summaries and topic breakdowns from recorded meetings.
- **Sales call analysis**: sentiment and entity detection across sales calls.
- **Contact center analytics**: topic detection and sentiment scoring at scale.
- **Medical documentation**: structured summaries from clinical audio.
- **Content repurposing**: chapters and summaries generated directly from transcripts.

---

## Ecosystem and Integrations

Speech Understanding add-ons apply to output from both [Speech-to-Text](/tech/assemblyai/assemblyai-speech-to-text) and [Universal-Streaming](/tech/assemblyai/assemblyai-streaming-speech-to-text), and can be combined with [Guardrails](/tech/assemblyai/assemblyai-guardrails) add-ons on the same request.

---

Read the [Speech Understanding docs](https://www.assemblyai.com/products/speech-understanding) to add these features to an existing transcription request.
61 changes: 61 additions & 0 deletions technologies/assemblyai/assemblyai-streaming-speech-to-text.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,61 @@
---
title: "Universal-Streaming"
author: "AssemblyAI"
authorUsername: "stevekimoi"
description: "Universal-Streaming is AssemblyAI's real-time speech-to-text API, delivering live transcription over WebSocket with the Universal-3.5 Pro Realtime model for voice agents and captioning."
---

# Universal-Streaming

Universal-Streaming is AssemblyAI's real-time speech-to-text product, delivering live transcription over a WebSocket connection with roughly 150ms latency. Its current model, Universal-3.5 Pro Realtime, is the successor to Universal-3 Pro Realtime and is built for voice agents, live captioning, and real-time agent-assist tools.

| General | |
|---------|--|
| **Developer** | AssemblyAI |
| **Type** | Real-time (streaming) speech-to-text API |
| **License** | Commercial API |
| **Documentation** | [assemblyai.com/docs](https://www.assemblyai.com/docs) |
| **GitHub** | [AssemblyAI/realtime-transcription-browser-js-example](https://github.com/AssemblyAI/realtime-transcription-browser-js-example) |

---

## Core Features

- **Low-latency streaming**: live audio is transcribed over WebSocket at approximately 150ms latency.
- **Universal-3.5 Pro Realtime model**: scores 4.1% WER on AssemblyAI's AA-WER Streaming benchmark, with roughly 0.4 seconds to the first final result.
- **Three latency/accuracy modes**: Balanced (default), Max Accuracy, and Min Latency.
- **Persistent conversation context**: context can be passed at the start of a call and after each agent turn without reconnecting the WebSocket.

Source: [Universal-3.5 Pro Realtime announcement](https://x.com/ArtificialAnlys/status/2074160133702402314), AssemblyAI docs hub.

---

## Pricing

| Model | Rate |
|-------|------|
| **Universal-3.5 Pro Realtime** | $0.45/hour ($7.50 per 1,000 minutes) |
| **Universal-Streaming** (English or Multilingual) | $0.15/hour |

Streaming is billed by WebSocket session duration, including idle time, rather than by audio duration. Streaming diarization adds $0.12/hour.

Source: [assemblyai.com/pricing](https://www.assemblyai.com/pricing)

---

## Tools and Resources

- **[Node.js/TypeScript SDK](https://github.com/AssemblyAI/assemblyai-node-sdk)**: official streaming client.
- **[Python SDK](https://github.com/AssemblyAI/assemblyai-python-sdk)**: official streaming client.
- **[Realtime browser example](https://github.com/AssemblyAI/realtime-transcription-browser-js-example)**: reference implementation for browser-based streaming.
- **[Streaming self-hosting stack](https://github.com/AssemblyAI/streaming-self-hosting-stack)**: Python stack for self-hosted streaming deployments.

---

## Ecosystem and Integrations

Universal-Streaming powers the listening layer behind AssemblyAI's [Voice Agent API](/tech/assemblyai/assemblyai-voice-agent-api), and is also available as a standalone streaming transcription endpoint for custom pipelines.

---

Get started with the [streaming docs](https://www.assemblyai.com/docs) and the $50 free credit available on signup.
57 changes: 57 additions & 0 deletions technologies/assemblyai/assemblyai-voice-agent-api.mdx
Original file line number Diff line number Diff line change
@@ -0,0 +1,57 @@
---
title: "Voice Agent API"
author: "AssemblyAI"
authorUsername: "stevekimoi"
description: "AssemblyAI's Voice Agent API is a fully managed speech-to-speech pipeline combining transcription, LLM reasoning, and text-to-speech over one WebSocket, with native LiveKit and Pipecat support."
---

# Voice Agent API

The [Voice Agent API](https://www.assemblyai.com/products/voice-agent-api) is AssemblyAI's fully managed speech-to-speech pipeline. It combines speech-to-text, LLM reasoning, text-to-speech, turn detection, interruption handling, and tool calling behind a single WebSocket connection, with an end-to-end latency of roughly one second.

| General | |
|---------|--|
| **Developer** | AssemblyAI |
| **Type** | Managed voice agent / speech-to-speech API |
| **License** | Commercial API |
| **Documentation** | [assemblyai.com/products/voice-agent-api](https://www.assemblyai.com/products/voice-agent-api) |
| **GitHub** | [AssemblyAI/voice-agent-starter-python](https://github.com/AssemblyAI/voice-agent-starter-python) |

---

## Core Features

- **Single-WebSocket pipeline**: transcription, LLM reasoning, and text-to-speech run behind one connection instead of separate services to wire together.
- **Turn detection and interruption handling**: manages conversational turn-taking so an agent can be interrupted mid-response.
- **Tool calling**: agents can call external tools/functions as part of a conversation.
- **PCI certification**: the pipeline is PCI-certified end-to-end, relevant for payment-related voice use cases.
- **Native framework plugins**: one-line integration with LiveKit (via `livekit-agents` 1.6+ and the `assemblyai` plugin) and Pipecat, both using Universal-3.5 Pro under the hood.
- **Telephony support**: works with Twilio and Telnyx out of the box.

---

## Pricing

The Voice Agent API is billed at a flat $4.50/hour ($0.075/minute) for the full pipeline, rather than pricing each component (STT, LLM, TTS) separately.

Source: [assemblyai.com/pricing](https://www.assemblyai.com/pricing)

---

## Tools and Resources

- **[Voice agent starter (Python)](https://github.com/AssemblyAI/voice-agent-starter-python)**: reference implementation to build from.
- **[Voice agent starter (JS)](https://github.com/AssemblyAI/voice-agent-starter-js)**: JavaScript/TypeScript starter template.
- **[LiveKit integration guide](https://www.assemblyai.com/blog/build-a-voice-agent-with-livekit-voice-agent-api)**: walkthrough for building a voice agent with LiveKit.

---

## Ecosystem and Integrations

- Native plugins for **LiveKit** and **Pipecat**.
- Out-of-the-box telephony support via **Twilio** and **Telnyx**.
- Built on AssemblyAI's [Universal-Streaming](/tech/assemblyai/assemblyai-streaming-speech-to-text) model for the listening layer.

---

Start with the [voice agent starter templates](https://github.com/AssemblyAI/voice-agent-starter-python) or the [LiveKit integration guide](https://www.assemblyai.com/blog/build-a-voice-agent-with-livekit-voice-agent-api).
Loading
Loading