Skip to content

Instantly share code, notes, and snippets.

@TosinAF
Last active April 2, 2026 13:44
Show Gist options
  • Select an option

  • Save TosinAF/fc259581c3c6d84c31603c40166e232f to your computer and use it in GitHub Desktop.

Select an option

Save TosinAF/fc259581c3c6d84c31603c40166e232f to your computer and use it in GitHub Desktop.
[ERD] Text-to-Speech for Harvey Responses — dual approach with Mixpanel tracking

[ERD] Text-to-Speech for Harvey Responses

Related Links


Summary

Add text-to-speech to Harvey so users can listen to assistant responses on Web and Mobile. Primary framing: accessibility — WCAG compliance for TTS could unlock enterprise deals with accessibility requirements (CAF Development Bank explicitly asked; Stein Adler wants expert report audio for reducing reading burden).

Two delivery approaches (full file + progressive chunks) will be spiked behind separate API routes, tested side-by-side, and the winning route productized.

What it does: User taps "play" on an assistant response. If the response contains complex formatting (tables, citations, multi-language content), the backend generates an audio-friendly version first. Audio plays with pause, resume, and +/- 15 sec skip controls. Original response stays on screen with a note: "Playing audio-optimized version."

What it doesn't do (V1):

  • Real-time TTS while the LLM is still generating (requires two-stream sync — out of scope, contingent on V1 usage)
  • Voice selection (single hardcoded default voice)
  • Text highlighting synced with audio (P2)
  • Scrub bar UI (P1 — mechanically works, just not in V1 UI)
  • Audio caching (evaluate after V1 usage data)

Model: ElevenLabs Flash v2.5 (eleven_flash_v2_5) — ~75ms latency, 32 languages, 40k char limit, 50% cheaper than Multilingual v2.

Voice: Single hardcoded default (TBD — pick a good ElevenLabs pre-built voice). Joey confirmed we could use a marketing-style voice but users won't be able to select in V1.

Vendor approval: ElevenLabs covered under the same sec/trust review as transcription (Scribe). Already in progress via #epd-sectrust-review (Anna, 2026-03-31).


Customer Feedback & Signal Landscape

Direct TTS (Text-to-Speech Output) Requests

# Customer Contact Use Case Source Date
1 Stein Adler Jake Convert large expert reports to spoken audio — reduce reading burden for attorneys Notion Feedback Jan 2025
2 Syensqo (via Notion) Read aloud for summaries while multitasking — "Please consider adding a read aloud function so that Harvey can read its response to the user. People like to listen and multitask while Harvey reads its summary" Notion Feedback Apr 2025
3 CAF Development Bank (via Gong, Spanish) Accessibility / WCAG compliance — audio/TTS for users with visual, motor, and neurological disabilities. Three official languages, compliance obligation. Gong (Spanish) Dec 2025
4 Paul Weiss Lauren Hudson Text-to-audio for professional development content — wants audio output from Harvey responses Gong Jan 2026
5 (unnamed partner) (via Slack) Reading docs aloud while driving — wants to listen to Harvey outputs on the go Slack (internal) Feb 2026
6 Paramount Paul Fagan Asked about speech output from Harvey Gong Feb 2026
7 — Anna Zhang (Harvey) Confirmed in a Gong call that TTS is being worked on — validates internal prioritization Gong Jan 2026

7 distinct signals spanning 3 clear demand drivers:

  1. Accessibility / compliance — CAF explicitly asked for WCAG-compliant audio. Could be a procurement blocker for regulated enterprises.
  2. Multitasking / mobile — Syensqo wants read-aloud while multitasking. Unnamed partner wants to listen while driving. Classic "eyes-free" use case.
  3. Long document consumption — Stein Adler wants expert reports as audio. Paul Weiss wants professional development content as audio. Reduces reading burden for dense legal material.

TTS has zero presence in the Enterpret Wisdom taxonomy — no theme exists for it. This is a genuinely novel feature request, not an underserved existing category.

Data source gaps: The CAF record doesn't surface via the Wisdom MCP search API (Spanish-language Gong transcript with accessibility framing). Several records (Syensqo, Paul Weiss, Paramount) were found via Enterpret web dashboard full-text search but don't surface cleanly via Cypher queries — the feedback is embedded in broader conversation transcripts where TTS is mentioned briefly.

Adjacent Signal: Voice Dictation (Input) — Validates Voice Interface Demand

Voice dictation is already shipping and actively used. In the last week alone, 23+ Gong records across 13 dedicated Wisdom themes:

Signal Customers Source
Praise for dictation Gramelis, Abrams Law, Plains GP, Crinetics, Arrise, Hachar Law, Saxena White, Lawson Lundell Gong
Dictation misrecognizes speech Morrison Mahoney, Philippi Prietocarrizosa Gong
Enable ubiquitous dictation access (web, not just mobile) Farella Braun, Universal Health, Morrison Mahoney, Kean Miller Gong
Improve live dictation controls Galileo Financial, Universal Health Gong
Voice dictation feature request Harneys (Outlook add-in), Mattos Filho (cited ChatGPT), BASF, Amereller (persistent sessions, undo/redo) Notion Feedback
"AI should reduce typing" Zahid Group (Matto Grassani explicitly said voice is key) Gong

Linear: FB-1930 — Real-time dictation for voice-to-prompt

Why this matters for TTS: Dictation (voice in) and TTS (voice out) are two sides of the same voice interface. The dictation signal validates that customers want voice interactions with Harvey. TTS completes the loop.

Adjacent Signal: Audio Transcription (Speech-to-Text) — Highest Volume

Audio transcription is an existing feature with 15+ customer enhancement requests — the largest cluster in the voice/audio space. Covered by Anna's PRD and Jin's ERD (linked above). Included here for context:

Customer Request
Godoy Cordoba Accuracy issues, 45-min delays, 2hr limit too short (hearings = 4hrs)
Conyers Court audio often >2hrs
Mayer Brown Bulk transcription ("thousands" of files, only 10 at a time)
Jones Day, Beccar Verela, Lightfoot Franklin, Arnall Golden Gregory Audio uploads in Workflow Builder
Dentons Auto-transcription for audio uploaded to Vault
Godoy, Mori Hamada, Repsol Language-specific quality (Colombian Spanish, Japanese, English-only output)
Chiomenti Transcription misses ~2-3 min of audio (bug)

9 Linear tickets: FB-3783, FB-3755, FB-2831, FB-2714, FB-3614, FB-2816, FB-1984, FB-1465, FB-480

Signal Summary

Category Volume Drivers Data Source Gap
TTS output (this feature) 7 records, 5 customers + 1 internal Accessibility, multitasking, long doc consumption Not in Wisdom taxonomy; several records not searchable via API
Voice dictation input 12+ customers Active use, quality issues, web access demand Wisdom has 13 themes; Notion Feedback not indexed
Audio transcription 15+ customers, 9 Linear tickets Accuracy, language, duration, bulk processing Covered by separate PRD/ERD

Justification for TTS despite modest signal:

  1. Accessibility compliance — could be a procurement blocker (CAF)
  2. Completes the voice interface loop (dictation in -> TTS out)
  3. Greenfield — no competitor has this for legal AI
  4. Low signal may reflect that nobody expects a legal AI tool to have it yet
  5. Mixpanel tracking from day 1 to validate demand once shipped
  6. Three distinct use case drivers gives confidence this isn't a one-off request

Requirements

Priority Requirement Source
P0 Pause / Resume Anna
P0 +/- 15 sec skip (this IS seeking) Anna, corrected
P0 Works on Web and iOS Team consensus
P0 Mixpanel events for usage tracking Tosin — validate usage before further investment
P0 Audio-friendly response transformation for complex content Tosin — needs shaping (see below)
P1 Scrub bar (drag to arbitrary position) Anna
P2 Text highlighting synced with audio Joey's idea, Anna confirmed P2

Non-Goals (V1)

  • Real-time LLM-to-speech streaming (Jin's sync concern is valid — not worth the complexity until V1 usage validated)
  • Voice selection / voice picker
  • Audio caching (evaluate after V1 usage data)

Audio-Friendly Response Transformation (Needs Shaping)

The Problem

Not all assistant responses are suitable for direct TTS. Responses may contain:

  • Tables — nonsensical when read aloud ("Header 1 pipe Header 2 pipe...")
  • Code blocks — meaningless as audio
  • Long bullet lists — lose structure in spoken form
  • Citations / footnotes — disrupt listening flow
  • Multi-language content — TTS voice may not handle language switching
  • Legal formatting — numbered sections, cross-references, defined terms

Current Thinking (Gut Instinct)

Backend auto-transforms the response text before sending to ElevenLabs — strip/simplify complex formatting without user interaction. Original response stays on screen; audio plays the transformed version with a small note: "Playing audio-optimized version."

What Needs Shaping

Before deciding the spike scope, we need to explore this problem space properly using the shaping skill. Key questions:

  1. Where is the boundary between "strip formatting" and "re-generate the response"? Stripping a table's pipe characters is different from asking the LLM to summarize the table as prose.
  2. Should the transformation be a simple text cleanup, an LLM re-prompt, or both? Simple cleanup is fast and cheap. LLM re-prompt produces better audio but adds latency and cost.
  3. What's the UX for responses that are fundamentally non-audio-friendly? A 50-row comparison table might not work as audio at all. Do we refuse, warn, or attempt anyway?
  4. Does this change the shape of the API? If we're re-prompting the LLM, the "text" input to the TTS route isn't the raw response anymore — it's a derived artifact.
  5. Can we automate the decision? User presses play → backend detects signals → auto-transforms → plays. Or does the user need to be in the loop?
  6. How does this interact with the response being streamed? If the response is still arriving, we can't analyze it for audio-friendliness yet.

Action: Run /shaping session before finalizing spike scope.


Architecture

Platform Order

  1. Spike: BE + Web — prove both routes work, compare feel, validate ElevenLabs quality/latency
  2. Productize: Web — ship winning route on Web first
  3. iOS — follows after Web is validated

All work happens in harveyai/app monorepo (mirrors harveyai/backend structure). Tosin builds all three (BE + Web + iOS).

Spike Success Criteria

The spike is done when:

  • Both routes work end-to-end (text in → audio plays on Web with controls)
  • Routes are compared on real devices for feel (latency, seeking responsiveness)
  • ElevenLabs Flash v2.5 quality and latency validated as acceptable
  • A route recommendation is made with evidence

Two API Routes (Built in Parallel, Tested Side-by-Side)

Route 1: POST /api/tts/convert

Full audio file. Backend buffers entire ElevenLabs response, serves as complete MP3 with Content-Length + Accept-Ranges: bytes.

Client → POST /api/tts/convert { text, voice_id? }
       ← 200 audio/mpeg (complete file, Content-Length set)
  • All seeking works immediately on both platforms
  • User waits for generation (~75ms/sentence with Flash v2.5, <1s for short responses per Jamie @ ElevenLabs)
  • Simplest client implementation

Route 2: POST /api/tts/stream

Progressive chunks. Backend streams ElevenLabs response as MP3 frames, client builds growing buffer.

Client → POST /api/tts/stream { text, voice_id? }
       ← 200 audio/mpeg (chunked transfer, frames arrive progressively)
  • Playback starts after ~1-2s of audio buffered
  • Seeking works within received range; forward skip beyond buffer queues until audio arrives
  • More complex client code (Web: MediaSource API, iOS: AVAudioPlayer reload with growing Data)

Shared Components

tts/
├── elevenlabs_client.py     # ElevenLabs API wrapper (auth, retry, error handling)
├── service.py               # TTSService — orchestrates generation, calls client
├── routes.py                # /api/tts/convert + /api/tts/stream endpoints
├── config.py                # Voice/model config, feature flags
├── transform.py             # Audio-friendly text transformation (scope TBD after shaping)
└── types.py                 # Request/response types

ElevenLabs Client

Wraps the ElevenLabs HTTP API:

  • Auth: API key from config/secrets
  • Endpoints used:
    • POST /v1/text-to-speech/{voice_id} — non-streaming, returns complete audio (Route 1)
    • POST /v1/text-to-speech/{voice_id}/stream — streaming, returns chunked audio (Route 2)
  • Model: eleven_flash_v2_5 (configurable via Statsig for future model swaps)
  • Retry: 1 retry on 5xx, no retry on 4xx
  • Timeout: 60s (generous — Flash v2.5 generates 40k chars in ~30s worst case)

Data Model

No new database tables for V1. Audio is ephemeral — generated on demand, not persisted.

Future consideration (post-V1 if caching is needed):

CREATE TABLE tts_audio_cache (
    id              UUID PRIMARY KEY DEFAULT uuid_v7(),
    workspace_id    UUID NOT NULL REFERENCES workspaces(id),
    text_hash       VARCHAR(64) NOT NULL,  -- SHA-256 of input text
    voice_id        VARCHAR(64) NOT NULL,
    model_id        VARCHAR(64) NOT NULL,
    audio_blob_url  TEXT NOT NULL,          -- Azure blob URL
    duration_ms     INTEGER,
    created_at      TIMESTAMP NOT NULL DEFAULT NOW(),
    updated_at      TIMESTAMP NOT NULL DEFAULT NOW(),
    deleted_at      TIMESTAMP,

    CONSTRAINT uq_tts_cache UNIQUE (workspace_id, text_hash, voice_id, model_id)
);

CREATE INDEX idx_tts_cache_workspace ON tts_audio_cache(workspace_id);

This is explicitly NOT in V1 scope. Including here so the design doesn't preclude it.


Analytics (Mixpanel)

Critical for V1 — we need to validate usage before investing in streaming complexity, caching, or real-time LLM-to-speech.

Events

Event Properties When
tts_playback_started route (convert/stream), text_length, response_id, platform (web/ios), model_id, voice_id, was_transformed User taps play, audio begins
tts_playback_completed route, duration_listened_ms, total_duration_ms, completion_pct, platform Audio finishes or user navigates away
tts_playback_paused route, position_ms, platform User pauses
tts_playback_resumed route, position_ms, platform User resumes
tts_skip_forward route, from_ms, to_ms, platform User skips +15s
tts_skip_backward route, from_ms, to_ms, platform User skips -15s
tts_generation_requested route, text_length, platform, has_complex_content Backend receives TTS request
tts_generation_completed route, text_length, generation_time_ms, audio_duration_ms, model_id, was_transformed ElevenLabs returns audio
tts_generation_failed route, text_length, error_type, model_id ElevenLabs call fails

Key Metrics to Track

Metric Why
Daily active TTS users Are people using this at all?
Completion rate (% of audio listened) Do users listen to the end or bail early?
Skip frequency How often do they use +/- 15s? Validates P0 priority of seeking.
Route A vs Route B preference During A/B test — which feels better?
Text length distribution Are they TTS-ing short or long responses? Informs caching + streaming decisions.
Transformation rate How often do responses need audio-friendly transformation?
Generation latency p50/p95 Is the wait acceptable?
Error rate ElevenLabs reliability

Observability (Datadog)

Metric Type Tags
tts.request.started counter route, platform
tts.request.completed counter route, platform, status
tts.request.duration_ms histogram route
tts.elevenlabs.latency_ms histogram model, endpoint
tts.elevenlabs.error counter model, error_type, status_code
tts.text_length histogram route
tts.transform.applied counter transform_type

Alerts

Alert Condition Severity
ElevenLabs error rate spike >5% over 5 min P2
ElevenLabs latency regression p95 >10s over 5 min P2
TTS request volume drop >50% drop vs previous day P3 (usage anomaly)

Feature Flag

Statsig gate: ENABLE_TTS_PLAYBACK

  • Controls whether TTS play button is shown in UI
  • Allows gradual rollout (internal -> pilot -> GA)
  • Separate from the backend routes (routes always available for testing, UI gate controls visibility)

Statsig Dynamic Config: TTS_CONFIG

{
    "model_id": "eleven_flash_v2_5",
    "voice_id": "<default-voice-id>",
    "max_text_length": 40000,
    "request_timeout_seconds": 60,
    "enable_stream_route": true,
    "enable_convert_route": true
}

Execution Plan

Pre-Spike: Shape the audio-friendly transformation problem

  • Run /shaping to explore the problem space
  • Output: decision on whether transformation is in spike scope, and if so, what approach

Phase 1: Spike (BE + Web)

  1. elevenlabs_client.py — API wrapper with auth + retry
  2. service.py — TTSService orchestration
  3. routes.py — both /api/tts/convert and /api/tts/stream
  4. Basic Web audio player (pause, resume, +/- 15s)
  5. Wire up both routes on Web
  6. Compare routes: latency, seeking feel, quality
  7. Spike output: route recommendation with evidence

Phase 2: Productize (BE + Web)

  1. Winning route hardened (error handling, rate limiting, timeouts)
  2. Audio-friendly transformation (scope from shaping session)
  3. Mixpanel events (both backend + client-side)
  4. Datadog metrics + alerts
  5. Feature flag gate
  6. Unit tests
  7. Dogfood -> Pilot -> GA

Phase 3: iOS

  1. Audio player UI (same controls as Web)
  2. Wire up winning route
  3. Mixpanel events
  4. Feature flag gate

Phase 4: Evaluate and Decide Next Investment

  • Analyze Mixpanel data: usage, completion rates, skip frequency
  • Decide whether to invest in: caching, real-time LLM-to-speech, text highlighting, scrub bar, voice selection

Rollout

Stage Scope Criteria to proceed
Dogfood Internal workspace Functional, no crashes, acceptable quality
Pilot 2-3 customer workspaces 1 week, error rate <1%, p95 latency <5s
GA All workspaces Stable 2+ weeks, Mixpanel shows sustained usage

Risks and Mitigations

Risk Mitigation
ElevenLabs API outage Graceful degradation — hide TTS button when generation fails. No fallback provider in V1.
Low adoption Mixpanel tracking from day 1. If DAU is low after 2 weeks of GA, deprioritize further investment.
iOS streaming seeking glitches Both routes spiked on Web first — if progressive chunks are glitchy, full file is the safe fallback.
Safari MediaSource + MP3 compat Route 1 works everywhere. Route 2 on Safari < 17.1 falls back to Route 1.
Cost at scale Flash v2.5 is 50% cheaper. Monitor via Datadog. Add caching (Phase 4) if costs are high.
Audio-friendly transformation quality Shape the problem first. Start with simple text cleanup; LLM re-prompt is a future enhancement if needed.

Security Considerations

  • Data in transit: Text sent to ElevenLabs over HTTPS. Covered under same vendor approval as transcription (Scribe) — already in progress.
  • No audio persistence in V1: Audio is ephemeral, generated per-request. No blob storage, no new data at rest.
  • API key management: ElevenLabs API key stored in 1Password / secrets manager, injected via env config. Never exposed to clients.
  • No new auth surface: TTS routes use existing auth middleware. User must be authenticated and have access to the assistant response they're requesting TTS for.
  • Rate limiting: Per-user rate limit on TTS requests to prevent abuse / cost runaway.

Open Questions

  1. Default voice ID — Which ElevenLabs pre-built voice? Need to listen to options and pick one.
  2. Safari < 17.1 support — Do we need it? If yes, Route 2 needs AAC/fMP4 instead of MP3 for MediaSource.
  3. Audio-friendly transformation scope — Pending shaping session. May or may not be in spike.
  4. harveyai/app monorepo structure — Where exactly does TTS backend code live? Mirrors harveyai/backend — need to confirm directory.
  5. Text sanitization baseline — Even without full transformation, do we strip markdown before sending raw text to ElevenLabs?
  6. Accessibility framing for launch — Should we position this as an accessibility feature in release comms? Could help with enterprise procurement.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment