Audio Reports
Configure, generate, play, and retrieve narrated intelligence reports
Audio Reports turn a generated intelligence report into an MP3 narration. The feature is available in active Pro and Elite workspaces and is currently labeled Alpha in the profile editor.
Audio is configured per profile. Enabling narration affects future report runs; it does not generate audio for reports that already exist.
Enable narration for a profile
Open the profile editor
Go to InfoOps → My Profiles, then create a profile or edit an existing one.
Open Advanced
Continue to the Advanced Configuration step and find Audio Narration.
Turn on audio
Enable the switch. Free and Lite workspaces cannot enable it; the backend also rejects an enabled audio configuration when the active workspace lacks the entitlement.
Choose the narration settings
Preview and select a voice, choose a generation speed, and optionally assign a synced pronunciation dictionary. If no voice is selected, Y2 uses its default professional narrator.
Save the profile
Audio is generated the next time that profile produces a non-degraded report.
Prop
Type
The voice library groups available voices into News & Broadcast, Professional, Narrative & Documentary, and Casual & Conversational. The stored catalog can include English and Spanish voices with masculine, feminine, or neutral labels. Availability is data-driven, so the exact voice list can change.
The current dictionary selector lists all dictionaries in the active scope; it does not filter them to the selected voice's language. Check the language badge yourself before assigning one.
What happens during generation
Y2 prepares a speech-specific version of the report before calling the text-to-speech provider:
- It removes the trailing references section, citation markers, HTML, and Markdown formatting.
- It normalizes headings, punctuation, common entities, numbers, currency, acronyms, and tickers for spoken output.
- It adds the assigned branding template's audio intro, before-report audio block, after-report audio block, and outro. When no value is configured, Y2 branding defaults apply.
- It applies the assigned pronunciation dictionary only when that dictionary has a provider ID
and a
syncedstatus. - It generates an MP3 with Cartesia Sonic-3 at 44.1 kHz and 128 kbps, then stores the file and its metadata in Convex.
Audio output is limited to 20 MB. Generation is also skipped when its estimated cost would push the report above a configured per-report budget.
Audio is an optional workflow step. A missing API key, provider error, oversize result, budget limit, or degraded report can leave the report without audio. Report delivery continues when the audio step returns no file.
The stored duration is an estimate based on the prepared text, not a measurement read from the finished MP3. The in-app player replaces it with the browser-reported media duration after the file loads.
Create a pronunciation dictionary
Pronunciation dictionaries are also an Alpha feature. Pro workspaces can store up to 10 and Elite workspaces up to 100. Each dictionary can contain up to 200 entries.
Open Audio
Go to InfoOps → Audio and select Create under Pronunciation Dictionaries.
Define the dictionary
Enter a name, optional description, and English or Spanish language. Add entries individually or use the CSV bulk-import control.
Add replacements
Each entry stores a source word, a pronunciation value, and an Alias or IPA label. The current provider sync sends entries as text-to-alias replacements, regardless of the stored label, so use a readable replacement string and test the generated result.
Save and verify sync
Creating or editing a dictionary starts an automatic provider sync. Confirm that its badge becomes Synced. If it becomes Error, inspect the displayed error and use the refresh button to retry after correcting the entries.
Assign it to a profile
Return to the profile's Advanced Configuration, enable Audio Narration, select the dictionary, and save.
| Constraint | Enforced limit |
|---|---|
| Dictionaries in Pro | 10 |
| Dictionaries in Elite | 100 |
| Entries per dictionary | 200 |
| Dictionary name | 50 characters |
| Description | 200 characters |
| Source word | 100 characters |
| Replacement | 300 characters |
Editing entries moves the dictionary back to pending before it is synced again. If a selected
dictionary is pending, errored, deleted, or outside the profile's tenant scope, Y2 does not apply
it during generation.
Customize spoken branding
The InfoOps → Audio page also exposes the intro and outro for each branding template in the
active workspace. Use {name} to insert the profile name.
| Spoken content | Limit | Position |
|---|---|---|
| Audio intro | 100 characters | Before all other narration |
| Before-report audio | 200 characters | After the intro and optional label |
| After-report audio | 200 characters | Before the outro and optional label |
| Audio outro | 200 characters | At the end |
The before- and after-report audio fields live in the full Branding Templates editor. A template must be explicitly assigned to the profile for its custom values to apply.
Play and deliver audio
When audio is available, the report page displays Listen. The bottom-sheet player supports play and pause, seeking, download, and playback rates of 0.75x, 1x, 1.25x, 1.5x, and 2x. Voice and duration metadata are displayed when available.
Delivery surfaces expose audio differently:
| Surface | Current behavior |
|---|---|
| In-app report | Shows the player only when the report has an audio URL |
| Shows an “Audio Narration Available” callout linking to the report page; the MP3 is not attached | |
| Webhook | Reports available or unavailable, an estimated duration, and an API audio link |
| Reports API | Returns audio metadata or redirects to the stored MP3 URL |
Audio files are removed when their owning report, profile, workspace, or account is deleted and can also be affected by storage cleanup. Do not treat a returned storage URL as a permanent archive.
Retrieve audio through the API
Use an API key with the reports:audio scope, or complete the endpoint's x402 flow. The path uses
the report's public API ID.
curl --request GET \
--url "https://api.y2.dev/api/v1/reports/REPORT_ID/audio" \
--header "Authorization: Bearer YOUR_API_KEY"{
"data": {
"url": "https://example.convex.cloud/api/storage/…",
"duration": 847,
"durationSeconds": 847,
"durationFormatted": "00:14:07",
"format": "mp3",
"mimeType": "audio/mpeg",
"fileSizeBytes": 13420800
}
}duration is retained for compatibility and is deprecated; use durationSeconds. Add
?redirect=true, or request an audio media type with the Accept header, for a 302 response to
the stored file URL.
The related GET /api/v1/reports/{reportId}/audio-text endpoint requires reports:read and
returns the speech-preprocessed text, character count, and estimated duration without requiring
that an MP3 already exists. See the generated Reports reference for
canonical schemas, errors, rate limits, and x402 pricing.
Troubleshoot audio
Next steps
Create or edit a profile
Configure the report workflow that produces narration.
Customize branding
Add intros, outros, and channel-specific static content.
Use the Reports API
Retrieve reports and their derived representations programmatically.
Review plans and limits
Compare audio access and pronunciation-dictionary quotas.