refactor(tts): consolidate TTS services under lib/tts with stable entrypoint (#551)
* refactor(tts): move service module into domain folder * refactor(tts): move summarization helpers into domain folder * refactor(tts): add module entrypoint exports * docs(tts): add module documentation * refactor(web): route TTS imports through tts entrypoint * docs(agents): map TTS module documentation
This commit is contained in:
@@ -0,0 +1,134 @@
|
||||
# TTS Module Documentation
|
||||
|
||||
## Purpose
|
||||
This module provides server-side Text-to-Speech services using OpenAI's TTS API, along with text summarization and sanitization utilities for preparing content for speech synthesis.
|
||||
|
||||
## Entrypoints and structure
|
||||
- `packages/web/server/lib/tts/index.js`: Public entrypoint imported by `packages/web/server/index.js`.
|
||||
- `packages/web/server/lib/tts/service.js`: TTS service implementation with OpenAI integration.
|
||||
- `packages/web/server/lib/tts/summarization.js`: Text summarization and sanitization utilities using opencode.ai zen API.
|
||||
|
||||
## Public exports
|
||||
|
||||
### TTS Service (from service.js)
|
||||
- `ttsService`: Singleton instance of TTSService class.
|
||||
- `TTSService`: TTS service class for OpenAI audio generation.
|
||||
- `TTS_VOICES`: Array of supported OpenAI voice identifiers.
|
||||
|
||||
### Summarization (from summarization.js)
|
||||
- `summarizeText({ text, threshold, maxLength, zenModel })`: Summarizes text for TTS output using opencode.ai zen API.
|
||||
- `sanitizeForTTS(text)`: Sanitizes text by removing markdown, URLs, file paths, and other non-speakable content.
|
||||
|
||||
## Constants
|
||||
|
||||
### Voice identifiers
|
||||
- `TTS_VOICES`: Array of supported OpenAI voices: `['alloy', 'ash', 'ballad', 'coral', 'echo', 'fable', 'nova', 'onyx', 'sage', 'shimmer', 'verse', 'marin', 'cedar']`.
|
||||
|
||||
### Summarization defaults
|
||||
- `SUMMARIZE_TIMEOUT_MS`: 30000 (30 seconds timeout for zen API requests).
|
||||
|
||||
### Default values
|
||||
- `summarizeText` defaults: `threshold` = 200, `maxLength` = 500, `zenModel` = 'gpt-5-nano'.
|
||||
- `generateSpeechStream` defaults: `voice` = 'coral', `model` = 'gpt-4o-mini-tts', `speed` = 1.0.
|
||||
- `generateSpeechBuffer` defaults: `voice` = 'coral', `model` = 'gpt-4o-mini-tts', `speed` = 1.0.
|
||||
|
||||
## TTSService methods
|
||||
|
||||
### `isAvailable()`
|
||||
Returns boolean indicating whether OpenAI API key is configured (checks environment variable `OPENAI_API_KEY` or OpenCode auth file).
|
||||
|
||||
### `generateSpeechStream(options)`
|
||||
Generates speech and returns as a web stream for direct streaming to clients.
|
||||
- Options: `text` (required), `voice`, `model`, `speed`, `instructions`, `apiKey`.
|
||||
- Returns: `{ stream: ReadableStream, contentType: 'audio/mpeg' }`.
|
||||
- Throws: Error if API key not configured or text is empty.
|
||||
|
||||
### `generateSpeechBuffer(options)`
|
||||
Generates speech and returns as Buffer for caching purposes.
|
||||
- Options: `text` (required), `voice`, `model`, `speed`, `instructions`.
|
||||
- Returns: Buffer containing MP3 audio data.
|
||||
- Throws: Error if API key not configured or text is empty.
|
||||
|
||||
## Response contracts
|
||||
|
||||
### `summarizeText`
|
||||
Returns object with:
|
||||
- `summary`: Sanitized summary text or original text (if not summarized).
|
||||
- `summarized`: Boolean indicating if summarization was performed.
|
||||
- `reason`: Optional string explaining why summarization was skipped (e.g., 'Text under threshold', 'Request timed out').
|
||||
- `originalLength`: Optional number for original text length.
|
||||
- `summaryLength`: Optional number for summarized text length.
|
||||
|
||||
### `sanitizeForTTS`
|
||||
Returns sanitized string with markdown, URLs, file paths, and special characters removed.
|
||||
|
||||
### `generateSpeechStream`
|
||||
Returns object with:
|
||||
- `stream`: ReadableStream of MP3 audio data.
|
||||
- `contentType`: Always 'audio/mpeg'.
|
||||
|
||||
### `generateSpeechBuffer`
|
||||
Returns Buffer containing MP3 audio data.
|
||||
|
||||
## API key resolution
|
||||
OpenAI API keys are resolved in order:
|
||||
1. Environment variable `OPENAI_API_KEY`.
|
||||
2. OpenCode auth file (`auth.openai`, `auth.codex`, or `auth.chatgpt`).
|
||||
3. Supports both string format (just token) and object format (with `access` or `token` fields).
|
||||
|
||||
## Usage in web server
|
||||
The TTS module is used by `packages/web/server/index.js` for:
|
||||
- Generating speech streams for client playback.
|
||||
- Generating speech buffers for caching.
|
||||
- Summarizing long messages before TTS synthesis.
|
||||
- Sanitizing text to remove non-speakable content.
|
||||
|
||||
The server-side TTS approach bypasses mobile Safari's audio context restrictions by generating audio on the server and streaming to clients.
|
||||
|
||||
## Notes for contributors
|
||||
|
||||
### Adding new TTS features
|
||||
1. Add new methods to `packages/web/server/lib/tts/service.js` TTSService class.
|
||||
2. Export public functions from `packages/web/server/lib/tts/index.js`.
|
||||
3. Follow existing patterns for API key resolution and error handling.
|
||||
4. Ensure all text is sanitized before TTS synthesis.
|
||||
5. Consider adding new voice options to `TTS_VOICES` constant.
|
||||
|
||||
### Text sanitization
|
||||
- Always call `sanitizeForTTS` on text before passing to TTS generation.
|
||||
- The sanitization removes markdown, code blocks, URLs, file paths, shell commands, and special characters.
|
||||
- This prevents the TTS from reading out technical formatting that sounds unnatural.
|
||||
|
||||
### Error handling
|
||||
- `generateSpeechStream` and `generateSpeechBuffer` throw descriptive errors for missing API keys or empty text.
|
||||
- `summarizeText` catches zen API errors and falls back to original text with `summarized: false`.
|
||||
- All errors are logged to console with `[TTSService]` or `[Summarize]` prefix.
|
||||
|
||||
### API key management
|
||||
- TTSService caches OpenAI client instance and recreates when API key changes.
|
||||
- API key changes are detected by comparing with `_lastApiKey` property.
|
||||
- This allows dynamic API key updates without server restart.
|
||||
|
||||
### Testing
|
||||
- Run `bun run type-check`, `bun run lint`, and `bun run build` before finalizing changes.
|
||||
- Test API key resolution with environment variable and auth file.
|
||||
- Test speech generation with various text lengths and voice options.
|
||||
- Test summarization behavior above and below threshold.
|
||||
- Test sanitization with markdown, URLs, and code blocks.
|
||||
- Verify streaming and buffer generation produce valid MP3 audio.
|
||||
|
||||
## Verification notes
|
||||
|
||||
### Manual verification
|
||||
1. Configure OpenAI API key via environment variable or OpenCode settings.
|
||||
2. Test `ttsService.isAvailable()` returns true.
|
||||
3. Call `ttsService.generateSpeechStream({ text: 'Hello world' })` and verify stream is returned.
|
||||
4. Call `ttsService.generateSpeechBuffer({ text: 'Hello world' })` and verify Buffer is returned.
|
||||
5. Test `summarizeText` with text above and below threshold.
|
||||
6. Test `sanitizeForTTS` with markdown, URLs, and code blocks.
|
||||
|
||||
### API endpoint verification
|
||||
1. Start web server and access TTS endpoint via client.
|
||||
2. Verify audio plays correctly in browser.
|
||||
3. Test on mobile Safari to verify bypass of audio context restrictions.
|
||||
4. Test with long messages to verify summarization is triggered.
|
||||
@@ -0,0 +1,16 @@
|
||||
/**
|
||||
* TTS Module Entry Point
|
||||
*
|
||||
* Public export surface for the Text-to-Speech domain module.
|
||||
*/
|
||||
|
||||
export {
|
||||
ttsService,
|
||||
TTSService,
|
||||
TTS_VOICES,
|
||||
} from './service.js';
|
||||
|
||||
export {
|
||||
summarizeText,
|
||||
sanitizeForTTS,
|
||||
} from './summarization.js';
|
||||
@@ -0,0 +1,162 @@
|
||||
/**
|
||||
* Server-side Text-to-Speech Service
|
||||
*
|
||||
* Uses OpenAI's TTS API to generate audio on the server and stream it to clients.
|
||||
* This bypasses mobile Safari's audio context restrictions.
|
||||
*/
|
||||
|
||||
import OpenAI from 'openai';
|
||||
import { readAuthFile } from '../opencode/auth.js';
|
||||
|
||||
// Voice options from OpenAI
|
||||
export const TTS_VOICES = [
|
||||
'alloy', 'ash', 'ballad', 'coral', 'echo', 'fable',
|
||||
'nova', 'onyx', 'sage', 'shimmer', 'verse', 'marin', 'cedar'
|
||||
];
|
||||
|
||||
function getOpenAIApiKey() {
|
||||
// First check environment variable
|
||||
const envKey = process.env.OPENAI_API_KEY;
|
||||
if (envKey) {
|
||||
return envKey;
|
||||
}
|
||||
|
||||
// Then check opencode auth file (same as usage tracker)
|
||||
try {
|
||||
const auth = readAuthFile();
|
||||
// Check for openai, codex, or chatgpt aliases
|
||||
const openaiAuth = auth.openai || auth.codex || auth.chatgpt;
|
||||
if (openaiAuth) {
|
||||
// Handle both string format (just the token) and object format
|
||||
if (typeof openaiAuth === 'string') {
|
||||
return openaiAuth;
|
||||
}
|
||||
// Try access token first (OAuth), then regular token
|
||||
if (openaiAuth.access) {
|
||||
return openaiAuth.access;
|
||||
}
|
||||
if (openaiAuth.token) {
|
||||
return openaiAuth.token;
|
||||
}
|
||||
}
|
||||
} catch (error) {
|
||||
console.warn('[TTSService] Failed to read auth file:', error.message);
|
||||
}
|
||||
|
||||
return null;
|
||||
}
|
||||
|
||||
class TTSService {
|
||||
constructor() {
|
||||
this._client = null;
|
||||
this._lastApiKey = null;
|
||||
}
|
||||
|
||||
_getClient() {
|
||||
const apiKey = getOpenAIApiKey();
|
||||
|
||||
// If API key changed or client doesn't exist, create new client
|
||||
if (apiKey && (!this._client || this._lastApiKey !== apiKey)) {
|
||||
this._client = new OpenAI({ apiKey });
|
||||
this._lastApiKey = apiKey;
|
||||
}
|
||||
|
||||
return this._client;
|
||||
}
|
||||
|
||||
isAvailable() {
|
||||
return this._getClient() !== null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Generate speech and return as a stream
|
||||
*/
|
||||
async generateSpeechStream(options) {
|
||||
const {
|
||||
text,
|
||||
voice = 'coral',
|
||||
model = 'gpt-4o-mini-tts',
|
||||
speed = 1.0,
|
||||
instructions,
|
||||
apiKey
|
||||
} = options;
|
||||
|
||||
// Use provided API key or fall back to configured key
|
||||
let client;
|
||||
if (apiKey) {
|
||||
client = new OpenAI({ apiKey });
|
||||
} else {
|
||||
client = this._getClient();
|
||||
}
|
||||
|
||||
if (!client) {
|
||||
throw new Error('OpenAI API key not configured. Set OPENAI_API_KEY environment variable, configure OpenAI in OpenCode, or provide an API key in settings.');
|
||||
}
|
||||
|
||||
if (!text.trim()) {
|
||||
throw new Error('Text is required for TTS');
|
||||
}
|
||||
|
||||
try {
|
||||
console.log('[TTSService] Generating speech with voice:', voice, 'model:', model);
|
||||
const response = await client.audio.speech.create({
|
||||
model,
|
||||
voice,
|
||||
input: text,
|
||||
speed,
|
||||
...(instructions && { instructions }),
|
||||
response_format: 'mp3',
|
||||
});
|
||||
|
||||
// Convert the response to a web stream
|
||||
const stream = response.body;
|
||||
|
||||
return {
|
||||
stream,
|
||||
contentType: 'audio/mpeg',
|
||||
};
|
||||
} catch (error) {
|
||||
console.error('[TTSService] Error generating speech:', error);
|
||||
throw new Error(`Failed to generate speech: ${error.message || 'Unknown error'}`);
|
||||
}
|
||||
}
|
||||
|
||||
/**
|
||||
* Generate speech and return as a buffer (for caching)
|
||||
*/
|
||||
async generateSpeechBuffer(options) {
|
||||
const client = this._getClient();
|
||||
if (!client) {
|
||||
throw new Error('OpenAI API key not configured. Set OPENAI_API_KEY environment variable or configure OpenAI in OpenCode.');
|
||||
}
|
||||
|
||||
const {
|
||||
text,
|
||||
voice = 'coral',
|
||||
model = 'gpt-4o-mini-tts',
|
||||
speed = 1.0,
|
||||
instructions
|
||||
} = options;
|
||||
|
||||
try {
|
||||
const response = await client.audio.speech.create({
|
||||
model,
|
||||
voice,
|
||||
input: text,
|
||||
speed,
|
||||
...(instructions && { instructions }),
|
||||
response_format: 'mp3',
|
||||
});
|
||||
|
||||
const arrayBuffer = await response.arrayBuffer();
|
||||
return Buffer.from(arrayBuffer);
|
||||
} catch (error) {
|
||||
console.error('[TTSService] Error generating speech buffer:', error);
|
||||
throw error;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Export singleton instance
|
||||
export const ttsService = new TTSService();
|
||||
export { TTSService };
|
||||
@@ -0,0 +1,171 @@
|
||||
/**
|
||||
* Text Summarization Service
|
||||
*
|
||||
* Uses the opencode.ai zen API with gpt-5-nano for fast, lightweight summarization.
|
||||
* Used by all TTS implementations (Browser, Say, OpenAI).
|
||||
*/
|
||||
|
||||
function buildSummarizationPrompt(maxLength) {
|
||||
return `You are a text summarizer for text-to-speech output. Create a concise, natural-sounding summary that captures the key points. Keep the summary under ${maxLength} characters.
|
||||
|
||||
CRITICAL INSTRUCTIONS:
|
||||
1. Output ONLY the final summary - no thinking, no reasoning, no explanations
|
||||
2. Do not show your work or thought process
|
||||
3. Do not use any special characters, markdown, code, URLs, file paths, or formatting
|
||||
4. Do not include phrases like "Here's a summary" or "In summary"
|
||||
5. Just provide clean, speakable text that can be read aloud
|
||||
6. Stay within the ${maxLength} character limit
|
||||
|
||||
Your response should be ready to speak immediately.`;
|
||||
}
|
||||
|
||||
const SUMMARIZE_TIMEOUT_MS = 30_000;
|
||||
|
||||
/**
|
||||
* Sanitize text for TTS output
|
||||
* Removes markdown, URLs, file paths, and other non-speakable content
|
||||
*/
|
||||
export function sanitizeForTTS(text) {
|
||||
if (!text || typeof text !== 'string') return '';
|
||||
|
||||
return text
|
||||
// Remove markdown formatting
|
||||
.replace(/[*_~`#]/g, '')
|
||||
// Remove code blocks
|
||||
.replace(/```[\s\S]*?```/g, '')
|
||||
.replace(/`[^`]*`/g, '')
|
||||
// Remove shell-like command patterns
|
||||
.replace(/^\s*[$#>]\s*/gm, '')
|
||||
// Remove common shell operators
|
||||
.replace(/[|&;<>]/g, ' ')
|
||||
// Remove backslashes (escape characters)
|
||||
.replace(/\\/g, '')
|
||||
// Remove brackets that might be interpreted specially
|
||||
.replace(/[[\]{}()]/g, '')
|
||||
// Remove quotes that might cause issues
|
||||
.replace(/["']/g, '')
|
||||
// Remove URLs
|
||||
.replace(/https?:\/\/[^\s]+/g, ' a link ')
|
||||
// Remove file paths
|
||||
.replace(/\/[\w\-./]+/g, '')
|
||||
// Collapse multiple spaces/newlines
|
||||
.replace(/\s+/g, ' ')
|
||||
.trim();
|
||||
}
|
||||
|
||||
/**
|
||||
* Extract text from zen API response
|
||||
*/
|
||||
function extractZenOutputText(data) {
|
||||
if (!data || typeof data !== 'object') return null;
|
||||
const output = data.output;
|
||||
if (!Array.isArray(output)) return null;
|
||||
|
||||
const messageItem = output.find(
|
||||
(item) => item && typeof item === 'object' && item.type === 'message'
|
||||
);
|
||||
if (!messageItem) return null;
|
||||
|
||||
const content = messageItem.content;
|
||||
if (!Array.isArray(content)) return null;
|
||||
|
||||
const textItem = content.find(
|
||||
(item) => item && typeof item === 'object' && item.type === 'output_text'
|
||||
);
|
||||
|
||||
const text = typeof textItem?.text === 'string' ? textItem.text.trim() : '';
|
||||
return text || null;
|
||||
}
|
||||
|
||||
/**
|
||||
* Summarize text using the opencode.ai zen API
|
||||
*
|
||||
* @param {Object} options
|
||||
* @param {string} options.text - The text to summarize
|
||||
* @param {number} options.threshold - Character threshold (don't summarize if under this length)
|
||||
* @param {number} options.maxLength - Maximum character length for the summary output (50-2000)
|
||||
* @param {string} [options.zenModel] - Override zen model (defaults to gpt-5-nano)
|
||||
* @returns {Promise<{summary: string, summarized: boolean, reason?: string}>}
|
||||
*/
|
||||
export async function summarizeText({
|
||||
text,
|
||||
threshold = 200,
|
||||
maxLength = 500,
|
||||
zenModel,
|
||||
}) {
|
||||
// Don't summarize if text is under threshold
|
||||
if (!text || text.length <= threshold) {
|
||||
return {
|
||||
summary: sanitizeForTTS(text || ''),
|
||||
summarized: false,
|
||||
reason: text ? 'Text under threshold' : 'No text provided',
|
||||
};
|
||||
}
|
||||
|
||||
const controller = new AbortController();
|
||||
const timer = setTimeout(() => controller.abort(), SUMMARIZE_TIMEOUT_MS);
|
||||
|
||||
try {
|
||||
const prompt = buildSummarizationPrompt(maxLength);
|
||||
|
||||
const response = await fetch('https://opencode.ai/zen/v1/responses', {
|
||||
method: 'POST',
|
||||
headers: { 'Content-Type': 'application/json' },
|
||||
body: JSON.stringify({
|
||||
model: zenModel || 'gpt-5-nano',
|
||||
input: [
|
||||
{ role: 'user', content: `${prompt}\n\nText to summarize:\n${text}` },
|
||||
],
|
||||
stream: false,
|
||||
reasoning: { effort: 'low' },
|
||||
}),
|
||||
signal: controller.signal,
|
||||
});
|
||||
|
||||
if (!response.ok) {
|
||||
const errorBody = await response.json().catch(() => ({}));
|
||||
console.error('[Summarize] zen API error:', response.status, errorBody);
|
||||
return {
|
||||
summary: sanitizeForTTS(text),
|
||||
summarized: false,
|
||||
reason: `zen API returned ${response.status}`,
|
||||
};
|
||||
}
|
||||
|
||||
const data = await response.json();
|
||||
const summary = extractZenOutputText(data);
|
||||
|
||||
if (summary) {
|
||||
const sanitized = sanitizeForTTS(summary);
|
||||
return {
|
||||
summary: sanitized,
|
||||
summarized: true,
|
||||
originalLength: text.length,
|
||||
summaryLength: sanitized.length,
|
||||
};
|
||||
}
|
||||
|
||||
return {
|
||||
summary: sanitizeForTTS(text),
|
||||
summarized: false,
|
||||
reason: 'No response from model',
|
||||
};
|
||||
} catch (error) {
|
||||
if (error.name === 'AbortError') {
|
||||
console.error('[Summarize] Request timed out');
|
||||
return {
|
||||
summary: sanitizeForTTS(text),
|
||||
summarized: false,
|
||||
reason: 'Request timed out',
|
||||
};
|
||||
}
|
||||
console.error('[Summarize] Error:', error);
|
||||
return {
|
||||
summary: sanitizeForTTS(text),
|
||||
summarized: false,
|
||||
reason: error.message,
|
||||
};
|
||||
} finally {
|
||||
clearTimeout(timer);
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user