Y2 Elite workspaces are rolling out for teams
Y2Y2Docs

Audio Reports

Configure, generate, play, and retrieve narrated intelligence reports

Audio Reports turn a generated intelligence report into an MP3 narration. The feature is available in active Pro and Elite workspaces and is currently labeled Alpha in the profile editor.

Audio is configured per profile. Enabling narration affects future report runs; it does not generate audio for reports that already exist.

Enable narration for a profile

Open the profile editor

Go to InfoOps → My Profiles, then create a profile or edit an existing one.

Open Advanced

Continue to the Advanced Configuration step and find Audio Narration.

Turn on audio

Enable the switch. Free and Lite workspaces cannot enable it; the backend also rejects an enabled audio configuration when the active workspace lacks the entitlement.

Choose the narration settings

Preview and select a voice, choose a generation speed, and optionally assign a synced pronunciation dictionary. If no voice is selected, Y2 uses its default professional narrator.

Save the profile

Audio is generated the next time that profile produces a non-degraded report.

Prop

Type

The voice library groups available voices into News & Broadcast, Professional, Narrative & Documentary, and Casual & Conversational. The stored catalog can include English and Spanish voices with masculine, feminine, or neutral labels. Availability is data-driven, so the exact voice list can change.

The current dictionary selector lists all dictionaries in the active scope; it does not filter them to the selected voice's language. Check the language badge yourself before assigning one.

What happens during generation

Y2 prepares a speech-specific version of the report before calling the text-to-speech provider:

  • It removes the trailing references section, citation markers, HTML, and Markdown formatting.
  • It normalizes headings, punctuation, common entities, numbers, currency, acronyms, and tickers for spoken output.
  • It adds the assigned branding template's audio intro, before-report audio block, after-report audio block, and outro. When no value is configured, Y2 branding defaults apply.
  • It applies the assigned pronunciation dictionary only when that dictionary has a provider ID and a synced status.
  • It generates an MP3 with Cartesia Sonic-3 at 44.1 kHz and 128 kbps, then stores the file and its metadata in Convex.

Audio output is limited to 20 MB. Generation is also skipped when its estimated cost would push the report above a configured per-report budget.

Audio is an optional workflow step. A missing API key, provider error, oversize result, budget limit, or degraded report can leave the report without audio. Report delivery continues when the audio step returns no file.

The stored duration is an estimate based on the prepared text, not a measurement read from the finished MP3. The in-app player replaces it with the browser-reported media duration after the file loads.

Create a pronunciation dictionary

Pronunciation dictionaries are also an Alpha feature. Pro workspaces can store up to 10 and Elite workspaces up to 100. Each dictionary can contain up to 200 entries.

Open Audio

Go to InfoOps → Audio and select Create under Pronunciation Dictionaries.

Define the dictionary

Enter a name, optional description, and English or Spanish language. Add entries individually or use the CSV bulk-import control.

Add replacements

Each entry stores a source word, a pronunciation value, and an Alias or IPA label. The current provider sync sends entries as text-to-alias replacements, regardless of the stored label, so use a readable replacement string and test the generated result.

Save and verify sync

Creating or editing a dictionary starts an automatic provider sync. Confirm that its badge becomes Synced. If it becomes Error, inspect the displayed error and use the refresh button to retry after correcting the entries.

Assign it to a profile

Return to the profile's Advanced Configuration, enable Audio Narration, select the dictionary, and save.

ConstraintEnforced limit
Dictionaries in Pro10
Dictionaries in Elite100
Entries per dictionary200
Dictionary name50 characters
Description200 characters
Source word100 characters
Replacement300 characters

Editing entries moves the dictionary back to pending before it is synced again. If a selected dictionary is pending, errored, deleted, or outside the profile's tenant scope, Y2 does not apply it during generation.

Customize spoken branding

The InfoOps → Audio page also exposes the intro and outro for each branding template in the active workspace. Use {name} to insert the profile name.

Spoken contentLimitPosition
Audio intro100 charactersBefore all other narration
Before-report audio200 charactersAfter the intro and optional label
After-report audio200 charactersBefore the outro and optional label
Audio outro200 charactersAt the end

The before- and after-report audio fields live in the full Branding Templates editor. A template must be explicitly assigned to the profile for its custom values to apply.

Play and deliver audio

When audio is available, the report page displays Listen. The bottom-sheet player supports play and pause, seeking, download, and playback rates of 0.75x, 1x, 1.25x, 1.5x, and 2x. Voice and duration metadata are displayed when available.

Delivery surfaces expose audio differently:

SurfaceCurrent behavior
In-app reportShows the player only when the report has an audio URL
EmailShows an “Audio Narration Available” callout linking to the report page; the MP3 is not attached
WebhookReports available or unavailable, an estimated duration, and an API audio link
Reports APIReturns audio metadata or redirects to the stored MP3 URL

Audio files are removed when their owning report, profile, workspace, or account is deleted and can also be affected by storage cleanup. Do not treat a returned storage URL as a permanent archive.

Retrieve audio through the API

Use an API key with the reports:audio scope, or complete the endpoint's x402 flow. The path uses the report's public API ID.

curl --request GET \
  --url "https://api.y2.dev/api/v1/reports/REPORT_ID/audio" \
  --header "Authorization: Bearer YOUR_API_KEY"
{
  "data": {
    "url": "https://example.convex.cloud/api/storage/…",
    "duration": 847,
    "durationSeconds": 847,
    "durationFormatted": "00:14:07",
    "format": "mp3",
    "mimeType": "audio/mpeg",
    "fileSizeBytes": 13420800
  }
}

duration is retained for compatibility and is deprecated; use durationSeconds. Add ?redirect=true, or request an audio media type with the Accept header, for a 302 response to the stored file URL.

The related GET /api/v1/reports/{reportId}/audio-text endpoint requires reports:read and returns the speech-preprocessed text, character count, and estimated duration without requiring that an MP3 already exists. See the generated Reports reference for canonical schemas, errors, rate limits, and x402 pricing.

Troubleshoot audio

Next steps