- Implementation Plan: https://gist.github.com/TosinAF/dfe34acbef8cba0b74d4ac6d7f52acd8
- Audio Seeking Decision Doc: https://gist.github.com/TosinAF/26654dcc2c84dea4808ca4472e00e2fe
- Reference ERD (Audio Transcription): https://www.notion.so/harveyai/ERD-Audio-Video-Transcription-330ac3fcdd7a81c9b347c5560bb8f1b0
- Linear: MOBILE-16 (voice reliability), MOBILE-59 (TTS on response)
- ElevenLabs channel: #ext-harvey-elevenlabs
- Voice eng channel: #voice-eng
Add text-to-speech to Harvey so users can listen to assistant responses on Web and Mobile. Primary framing: accessibility — WCAG compliance for TTS could unlock enterprise deals with accessibility requirements (CAF Development Bank explicitly asked; Stein Adler wants expert report audio for reducing reading burden).
Two delivery approaches (full file + progressive chunks) will be spiked behind separate API routes, tested side-by-side, and the winning route productized.
What it does: User taps "play" on an assistant response. If the response contains complex formatting (tables, citations, multi-language content), the backend generates an audio-friendly version first. Audio plays with pause, resume, and +/- 15 sec skip controls. Original response stays on screen with a note: "Playing audio-optimized version."
What it doesn't do (V1):
- Real-time TTS while the LLM is still generating (requires two-stream sync — out of scope, contingent on V1 usage)
- Voice selection (single hardcoded default voice)
- Text highlighting synced with audio (P2)
- Scrub bar UI (P1 — mechanically works, just not in V1 UI)
- Audio caching (evaluate after V1 usage data)
Model: ElevenLabs Flash v2.5 (eleven_flash_v2_5) — ~75ms latency, 32 languages, 40k char limit, 50% cheaper than Multilingual v2.
Voice: Single hardcoded default (TBD — pick a good ElevenLabs pre-built voice). Joey confirmed we could use a marketing-style voice but users won't be able to select in V1.
Vendor approval: ElevenLabs covered under the same sec/trust review as transcription (Scribe). Already in progress via #epd-sectrust-review (Anna, 2026-03-31).
| # | Customer | Contact | Use Case | Source | Date |
|---|---|---|---|---|---|
| 1 | Stein Adler | Jake | Convert large expert reports to spoken audio — reduce reading burden for attorneys | Notion Feedback | Jan 2025 |
| 2 | Syensqo | (via Notion) | Read aloud for summaries while multitasking — "Please consider adding a read aloud function so that Harvey can read its response to the user. People like to listen and multitask while Harvey reads its summary" | Notion Feedback | Apr 2025 |
| 3 | CAF Development Bank | (via Gong, Spanish) | Accessibility / WCAG compliance — audio/TTS for users with visual, motor, and neurological disabilities. Three official languages, compliance obligation. | Gong (Spanish) | Dec 2025 |
| 4 | Paul Weiss | Lauren Hudson | Text-to-audio for professional development content — wants audio output from Harvey responses | Gong | Jan 2026 |
| 5 | (unnamed partner) | (via Slack) | Reading docs aloud while driving — wants to listen to Harvey outputs on the go | Slack (internal) | Feb 2026 |
| 6 | Paramount | Paul Fagan | Asked about speech output from Harvey | Gong | Feb 2026 |
| 7 | — | Anna Zhang (Harvey) | Confirmed in a Gong call that TTS is being worked on — validates internal prioritization | Gong | Jan 2026 |
7 distinct signals spanning 3 clear demand drivers:
- Accessibility / compliance — CAF explicitly asked for WCAG-compliant audio. Could be a procurement blocker for regulated enterprises.
- Multitasking / mobile — Syensqo wants read-aloud while multitasking. Unnamed partner wants to listen while driving. Classic "eyes-free" use case.
- Long document consumption — Stein Adler wants expert reports as audio. Paul Weiss wants professional development content as audio. Reduces reading burden for dense legal material.
TTS has zero presence in the Enterpret Wisdom taxonomy — no theme exists for it. This is a genuinely novel feature request, not an underserved existing category.
Data source gaps: The CAF record doesn't surface via the Wisdom MCP search API (Spanish-language Gong transcript with accessibility framing). Several records (Syensqo, Paul Weiss, Paramount) were found via Enterpret web dashboard full-text search but don't surface cleanly via Cypher queries — the feedback is embedded in broader conversation transcripts where TTS is mentioned briefly.
Voice dictation is already shipping and actively used. In the last week alone, 23+ Gong records across 13 dedicated Wisdom themes:
| Signal | Customers | Source |
|---|---|---|
| Praise for dictation | Gramelis, Abrams Law, Plains GP, Crinetics, Arrise, Hachar Law, Saxena White, Lawson Lundell | Gong |
| Dictation misrecognizes speech | Morrison Mahoney, Philippi Prietocarrizosa | Gong |
| Enable ubiquitous dictation access (web, not just mobile) | Farella Braun, Universal Health, Morrison Mahoney, Kean Miller | Gong |
| Improve live dictation controls | Galileo Financial, Universal Health | Gong |
| Voice dictation feature request | Harneys (Outlook add-in), Mattos Filho (cited ChatGPT), BASF, Amereller (persistent sessions, undo/redo) | Notion Feedback |
| "AI should reduce typing" | Zahid Group (Matto Grassani explicitly said voice is key) | Gong |
Linear: FB-1930 — Real-time dictation for voice-to-prompt
Why this matters for TTS: Dictation (voice in) and TTS (voice out) are two sides of the same voice interface. The dictation signal validates that customers want voice interactions with Harvey. TTS completes the loop.
Audio transcription is an existing feature with 15+ customer enhancement requests — the largest cluster in the voice/audio space. Covered by Anna's PRD and Jin's ERD (linked above). Included here for context:
| Customer | Request |
|---|---|
| Godoy Cordoba | Accuracy issues, 45-min delays, 2hr limit too short (hearings = 4hrs) |
| Conyers | Court audio often >2hrs |
| Mayer Brown | Bulk transcription ("thousands" of files, only 10 at a time) |
| Jones Day, Beccar Verela, Lightfoot Franklin, Arnall Golden Gregory | Audio uploads in Workflow Builder |
| Dentons | Auto-transcription for audio uploaded to Vault |
| Godoy, Mori Hamada, Repsol | Language-specific quality (Colombian Spanish, Japanese, English-only output) |
| Chiomenti | Transcription misses ~2-3 min of audio (bug) |
9 Linear tickets: FB-3783, FB-3755, FB-2831, FB-2714, FB-3614, FB-2816, FB-1984, FB-1465, FB-480
| Category | Volume | Drivers | Data Source Gap |
|---|---|---|---|
| TTS output (this feature) | 7 records, 5 customers + 1 internal | Accessibility, multitasking, long doc consumption | Not in Wisdom taxonomy; several records not searchable via API |
| Voice dictation input | 12+ customers | Active use, quality issues, web access demand | Wisdom has 13 themes; Notion Feedback not indexed |
| Audio transcription | 15+ customers, 9 Linear tickets | Accuracy, language, duration, bulk processing | Covered by separate PRD/ERD |
Justification for TTS despite modest signal:
- Accessibility compliance — could be a procurement blocker (CAF)
- Completes the voice interface loop (dictation in -> TTS out)
- Greenfield — no competitor has this for legal AI
- Low signal may reflect that nobody expects a legal AI tool to have it yet
- Mixpanel tracking from day 1 to validate demand once shipped
- Three distinct use case drivers gives confidence this isn't a one-off request
| Priority | Requirement | Source |
|---|---|---|
| P0 | Pause / Resume | Anna |
| P0 | +/- 15 sec skip (this IS seeking) | Anna, corrected |
| P0 | Works on Web and iOS | Team consensus |
| P0 | Mixpanel events for usage tracking | Tosin — validate usage before further investment |
| P0 | Audio-friendly response transformation for complex content | Tosin — needs shaping (see below) |
| P1 | Scrub bar (drag to arbitrary position) | Anna |
| P2 | Text highlighting synced with audio | Joey's idea, Anna confirmed P2 |
- Real-time LLM-to-speech streaming (Jin's sync concern is valid — not worth the complexity until V1 usage validated)
- Voice selection / voice picker
- Audio caching (evaluate after V1 usage data)
Not all assistant responses are suitable for direct TTS. Responses may contain:
- Tables — nonsensical when read aloud ("Header 1 pipe Header 2 pipe...")
- Code blocks — meaningless as audio
- Long bullet lists — lose structure in spoken form
- Citations / footnotes — disrupt listening flow
- Multi-language content — TTS voice may not handle language switching
- Legal formatting — numbered sections, cross-references, defined terms
Backend auto-transforms the response text before sending to ElevenLabs — strip/simplify complex formatting without user interaction. Original response stays on screen; audio plays the transformed version with a small note: "Playing audio-optimized version."
Before deciding the spike scope, we need to explore this problem space properly using the shaping skill. Key questions:
- Where is the boundary between "strip formatting" and "re-generate the response"? Stripping a table's pipe characters is different from asking the LLM to summarize the table as prose.
- Should the transformation be a simple text cleanup, an LLM re-prompt, or both? Simple cleanup is fast and cheap. LLM re-prompt produces better audio but adds latency and cost.
- What's the UX for responses that are fundamentally non-audio-friendly? A 50-row comparison table might not work as audio at all. Do we refuse, warn, or attempt anyway?
- Does this change the shape of the API? If we're re-prompting the LLM, the "text" input to the TTS route isn't the raw response anymore — it's a derived artifact.
- Can we automate the decision? User presses play → backend detects signals → auto-transforms → plays. Or does the user need to be in the loop?
- How does this interact with the response being streamed? If the response is still arriving, we can't analyze it for audio-friendliness yet.
Action: Run /shaping session before finalizing spike scope.
- Spike: BE + Web — prove both routes work, compare feel, validate ElevenLabs quality/latency
- Productize: Web — ship winning route on Web first
- iOS — follows after Web is validated
All work happens in harveyai/app monorepo (mirrors harveyai/backend structure). Tosin builds all three (BE + Web + iOS).
The spike is done when:
- Both routes work end-to-end (text in → audio plays on Web with controls)
- Routes are compared on real devices for feel (latency, seeking responsiveness)
- ElevenLabs Flash v2.5 quality and latency validated as acceptable
- A route recommendation is made with evidence
Full audio file. Backend buffers entire ElevenLabs response, serves as complete MP3 with Content-Length + Accept-Ranges: bytes.
Client → POST /api/tts/convert { text, voice_id? }
← 200 audio/mpeg (complete file, Content-Length set)
- All seeking works immediately on both platforms
- User waits for generation (~75ms/sentence with Flash v2.5, <1s for short responses per Jamie @ ElevenLabs)
- Simplest client implementation
Progressive chunks. Backend streams ElevenLabs response as MP3 frames, client builds growing buffer.
Client → POST /api/tts/stream { text, voice_id? }
← 200 audio/mpeg (chunked transfer, frames arrive progressively)
- Playback starts after ~1-2s of audio buffered
- Seeking works within received range; forward skip beyond buffer queues until audio arrives
- More complex client code (Web:
MediaSourceAPI, iOS:AVAudioPlayerreload with growingData)
tts/
├── elevenlabs_client.py # ElevenLabs API wrapper (auth, retry, error handling)
├── service.py # TTSService — orchestrates generation, calls client
├── routes.py # /api/tts/convert + /api/tts/stream endpoints
├── config.py # Voice/model config, feature flags
├── transform.py # Audio-friendly text transformation (scope TBD after shaping)
└── types.py # Request/response types
Wraps the ElevenLabs HTTP API:
- Auth: API key from config/secrets
- Endpoints used:
POST /v1/text-to-speech/{voice_id}— non-streaming, returns complete audio (Route 1)POST /v1/text-to-speech/{voice_id}/stream— streaming, returns chunked audio (Route 2)
- Model:
eleven_flash_v2_5(configurable via Statsig for future model swaps) - Retry: 1 retry on 5xx, no retry on 4xx
- Timeout: 60s (generous — Flash v2.5 generates 40k chars in ~30s worst case)
No new database tables for V1. Audio is ephemeral — generated on demand, not persisted.
CREATE TABLE tts_audio_cache (
id UUID PRIMARY KEY DEFAULT uuid_v7(),
workspace_id UUID NOT NULL REFERENCES workspaces(id),
text_hash VARCHAR(64) NOT NULL, -- SHA-256 of input text
voice_id VARCHAR(64) NOT NULL,
model_id VARCHAR(64) NOT NULL,
audio_blob_url TEXT NOT NULL, -- Azure blob URL
duration_ms INTEGER,
created_at TIMESTAMP NOT NULL DEFAULT NOW(),
updated_at TIMESTAMP NOT NULL DEFAULT NOW(),
deleted_at TIMESTAMP,
CONSTRAINT uq_tts_cache UNIQUE (workspace_id, text_hash, voice_id, model_id)
);
CREATE INDEX idx_tts_cache_workspace ON tts_audio_cache(workspace_id);This is explicitly NOT in V1 scope. Including here so the design doesn't preclude it.
Critical for V1 — we need to validate usage before investing in streaming complexity, caching, or real-time LLM-to-speech.
| Event | Properties | When |
|---|---|---|
tts_playback_started |
route (convert/stream), text_length, response_id, platform (web/ios), model_id, voice_id, was_transformed |
User taps play, audio begins |
tts_playback_completed |
route, duration_listened_ms, total_duration_ms, completion_pct, platform |
Audio finishes or user navigates away |
tts_playback_paused |
route, position_ms, platform |
User pauses |
tts_playback_resumed |
route, position_ms, platform |
User resumes |
tts_skip_forward |
route, from_ms, to_ms, platform |
User skips +15s |
tts_skip_backward |
route, from_ms, to_ms, platform |
User skips -15s |
tts_generation_requested |
route, text_length, platform, has_complex_content |
Backend receives TTS request |
tts_generation_completed |
route, text_length, generation_time_ms, audio_duration_ms, model_id, was_transformed |
ElevenLabs returns audio |
tts_generation_failed |
route, text_length, error_type, model_id |
ElevenLabs call fails |
| Metric | Why |
|---|---|
| Daily active TTS users | Are people using this at all? |
| Completion rate (% of audio listened) | Do users listen to the end or bail early? |
| Skip frequency | How often do they use +/- 15s? Validates P0 priority of seeking. |
| Route A vs Route B preference | During A/B test — which feels better? |
| Text length distribution | Are they TTS-ing short or long responses? Informs caching + streaming decisions. |
| Transformation rate | How often do responses need audio-friendly transformation? |
| Generation latency p50/p95 | Is the wait acceptable? |
| Error rate | ElevenLabs reliability |
| Metric | Type | Tags |
|---|---|---|
tts.request.started |
counter | route, platform |
tts.request.completed |
counter | route, platform, status |
tts.request.duration_ms |
histogram | route |
tts.elevenlabs.latency_ms |
histogram | model, endpoint |
tts.elevenlabs.error |
counter | model, error_type, status_code |
tts.text_length |
histogram | route |
tts.transform.applied |
counter | transform_type |
| Alert | Condition | Severity |
|---|---|---|
| ElevenLabs error rate spike | >5% over 5 min | P2 |
| ElevenLabs latency regression | p95 >10s over 5 min | P2 |
| TTS request volume drop | >50% drop vs previous day | P3 (usage anomaly) |
Statsig gate: ENABLE_TTS_PLAYBACK
- Controls whether TTS play button is shown in UI
- Allows gradual rollout (internal -> pilot -> GA)
- Separate from the backend routes (routes always available for testing, UI gate controls visibility)
{
"model_id": "eleven_flash_v2_5",
"voice_id": "<default-voice-id>",
"max_text_length": 40000,
"request_timeout_seconds": 60,
"enable_stream_route": true,
"enable_convert_route": true
}- Run
/shapingto explore the problem space - Output: decision on whether transformation is in spike scope, and if so, what approach
elevenlabs_client.py— API wrapper with auth + retryservice.py— TTSService orchestrationroutes.py— both/api/tts/convertand/api/tts/stream- Basic Web audio player (pause, resume, +/- 15s)
- Wire up both routes on Web
- Compare routes: latency, seeking feel, quality
- Spike output: route recommendation with evidence
- Winning route hardened (error handling, rate limiting, timeouts)
- Audio-friendly transformation (scope from shaping session)
- Mixpanel events (both backend + client-side)
- Datadog metrics + alerts
- Feature flag gate
- Unit tests
- Dogfood -> Pilot -> GA
- Audio player UI (same controls as Web)
- Wire up winning route
- Mixpanel events
- Feature flag gate
- Analyze Mixpanel data: usage, completion rates, skip frequency
- Decide whether to invest in: caching, real-time LLM-to-speech, text highlighting, scrub bar, voice selection
| Stage | Scope | Criteria to proceed |
|---|---|---|
| Dogfood | Internal workspace | Functional, no crashes, acceptable quality |
| Pilot | 2-3 customer workspaces | 1 week, error rate <1%, p95 latency <5s |
| GA | All workspaces | Stable 2+ weeks, Mixpanel shows sustained usage |
| Risk | Mitigation |
|---|---|
| ElevenLabs API outage | Graceful degradation — hide TTS button when generation fails. No fallback provider in V1. |
| Low adoption | Mixpanel tracking from day 1. If DAU is low after 2 weeks of GA, deprioritize further investment. |
| iOS streaming seeking glitches | Both routes spiked on Web first — if progressive chunks are glitchy, full file is the safe fallback. |
Safari MediaSource + MP3 compat |
Route 1 works everywhere. Route 2 on Safari < 17.1 falls back to Route 1. |
| Cost at scale | Flash v2.5 is 50% cheaper. Monitor via Datadog. Add caching (Phase 4) if costs are high. |
| Audio-friendly transformation quality | Shape the problem first. Start with simple text cleanup; LLM re-prompt is a future enhancement if needed. |
- Data in transit: Text sent to ElevenLabs over HTTPS. Covered under same vendor approval as transcription (Scribe) — already in progress.
- No audio persistence in V1: Audio is ephemeral, generated per-request. No blob storage, no new data at rest.
- API key management: ElevenLabs API key stored in 1Password / secrets manager, injected via env config. Never exposed to clients.
- No new auth surface: TTS routes use existing auth middleware. User must be authenticated and have access to the assistant response they're requesting TTS for.
- Rate limiting: Per-user rate limit on TTS requests to prevent abuse / cost runaway.
- Default voice ID — Which ElevenLabs pre-built voice? Need to listen to options and pick one.
- Safari < 17.1 support — Do we need it? If yes, Route 2 needs AAC/fMP4 instead of MP3 for
MediaSource. - Audio-friendly transformation scope — Pending shaping session. May or may not be in spike.
harveyai/appmonorepo structure — Where exactly does TTS backend code live? Mirrorsharveyai/backend— need to confirm directory.- Text sanitization baseline — Even without full transformation, do we strip markdown before sending raw text to ElevenLabs?
- Accessibility framing for launch — Should we position this as an accessibility feature in release comms? Could help with enterprise procurement.