Skip to content

Speak for developers

Every word in sync.
Inside your workflow.

Turn audio and video into timed transcripts, styled captions, and platform-ready exports through the API or MCP.

  • Word timestamps
  • Async jobs
  • Platform presets

API, SDKs, and MCP are in beta and open to every account.

workflow.shBeta
# 1. Create a project from a media URL
curl -X POST https://capxion.me/api/v1/projects \
  -H "Authorization: Bearer $SPEAK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "source_url": "https://example.com/interview.mp4",
        "kind": "video", "language": "auto" }'

# 2. Fetch the word-timed transcript once it's ready
curl https://capxion.me/api/v1/projects/$PROJECT_ID/transcript \
  -H "Authorization: Bearer $SPEAK_API_KEY"

# 3. Export for two platforms with burned-in captions
curl -X POST https://capxion.me/api/v1/projects/$PROJECT_ID/exports \
  -H "Authorization: Bearer $SPEAK_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{ "presets": ["youtube", "instagram_reels"],
        "burn_captions": true }'

transcript.json

{ "text": "Every",
  "start": 0.42,
  "end": 0.71 }

captions.vtt

00:00.420 --> 00:01.580
Every word counts

video.mp4

A podcast host speaking into a microphone

From media to finished assets.

A simple workflow from upload to export.

  1. 1

    Create a project

    Upload or reference your recording.

    POST /api/v1/projects
    { "source_url": "https://…/talk.mp4",
      "kind": "speech" }
  2. 2

    Transcribe and refine

    Review words and timing before export.

    PUT /api/v1/projects/{id}/transcript
    { "lines": [{ "words": [{ "text": "hello",
      "start": 0.42, "end": 0.71 }] }] }
  3. 3

    Render and retrieve

    Track jobs and download completed files.

    POST /api/v1/projects/{id}/exports
    { "presets": ["tiktok"],
      "range_start": 12, "range_end": 42 }

Two ways to build with Speak.

Same powerful transcription and captioning engine. Choose the integration that fits your workflow.

API

Bring caption workflows into your product.

  • Projects and media
  • Timed transcripts
  • Caption styles
  • Export jobs
  • Webhooks
{
  "id": "3f0c7a52-8d7e-4a51-9b1e-2f1c0d6e8a41",
  "title": "Interview",
  "status": "ready",
  "kind": "video",
  "source": "video",
  "language": "en",
  "duration": 184.6,
  "format": "9:16",
  "style": { "template": "clean", "placement": "bottom" },
  "error": null,
  "job_id": "c5d1e9f0-1a2b-4c3d-8e9f-0a1b2c3d4e5f",
  "created_at": "2026-10-10T14:01:07Z"
}
Explore the API

MCP

Give your AI agent a captioning toolkit.

Caption this interview for Reels and YouTube.

Speak tools
  • Create project
  • Transcribe media
  • Generate exports
  • Retrieve results
Explore MCP
API resources and SDKs

Use typed clients for authentication, retries, and polling. Projects hold your media; transcripts contain word timings; exports create platform files; webhooks report job changes.

TypeScriptExample
npm install @speak/sdk
PythonExample
pip install speak-sdk
Read the SDK and MCP setup guide

Keep long-running work moving.

Submit work, track progress, and retrieve the result when it’s ready.

Completion event · example

{
  "id": "6c0f2b9e-4d1a-4f3e-9a57-0e8b1c2d3f40",
  "type": "export.completed",
  "created_at": "2026-10-10T14:03:22Z",
  "data": {
    "project_id": "3f0c7a52-8d7e-4a51-9b1e-2f1c0d6e8a41",
    "export_id": "b81e2d4c-77a0-4c0f-a6f3-5d9e1c2b3a70",
    "status": "completed",
    "error": null
  }
}

Built around the assets you need.

Transcripts, captions, and exports formatted for real-world use.

Word-level transcripts

Get time-aligned transcripts with word-level timestamps.

{
  "lines": [{
    "words": [{
      "text": "Every",
      "start": 0.42,
      "end": 0.71
    }, …]
  }]
}

Styled caption videos

Generate beautifully styled captions, ready to share.

A podcast host speaking into a microphone

Every idea
starts with a conversation.

Subtitle files

Export SRT and VTT subtitles for easy use anywhere.

1
00:00:00,420 --> 00:00:01,580
Every word counts

2
00:00:01,900 --> 00:00:03,240
so say what you mean

Platform-ready exports

Render for the formats your audience is on.

16:99:161:1
WebVTT sample
WEBVTT

00:00:00.420 --> 00:00:01.580
Every word counts

00:00:01.900 --> 00:00:03.240
so say what you mean

Start with one recording.

Get from idea to finished captions in minutes.

  1. 1

    Create your account

    Set up your Speak account.

  2. 2

    Set up developer access

    Get access to the API or MCP.

  3. 3

    Run your first workflow

    Transcribe and export a real file.

Get started

Build captions into what you’re building.

Turn audio and video into transcripts, captions, and exports through the API or MCP.