Owner Profile
0xCEB77b01...79b5B133Text-to-speech — short (≤2k)
x402Natural speech, instantly - English (6 voices) and Korean. Kokoro + MeloTTS engines; wav or mp3 output; voice and speed control per request; up to 2,000 chars (Korean 1,200). Failed calls are never charged. For longer text, use Text-to-speech — narration.
Speech-to-text — clips (≤3 min)
x402Accurate transcription in seconds - 99 languages, auto-detected. Powered by Whisper. Send audio as base64 or just pass a public audio_url (wav/mp3/m4a/ogg/webm, up to 3 minutes); returns the transcript with language confidence and optional timestamps. Failed calls are never charged. For audio over 3 minutes, use Speech-to-text — long-form.
Speech-to-text — long-form (≤2h)
x402Transcribe long-form audio - podcasts, meetings, lectures up to 2 hours. Submit a public audio_url, get a jobId instantly, poll /paid/jobs/{jobId} for the transcript (free). Usage-based billing: $0.003 per minute of audio, authorized up front at a $0.36 ceiling and settled at the measured length. Failed jobs are never charged. For clips under 3 minutes, Speech-to-text — clips is instant and fixed-price.
Text-to-speech — narration (50k)
x402Narrate long-form text - articles, newsletters, audiobook chapters up to 50,000 characters, English (6 voices) and Korean. Submit text, get a jobId instantly, poll /paid/jobs/{jobId} and download the mp3 from its audioUrl (free). Usage-based billing: $0.003 per minute of generated audio, authorized at a $0.36 ceiling and settled at the measured length. Failed jobs are never charged. For short snippets under 2,000 chars, Text-to-speech — short returns audio instantly.