Home
AI Video Translation AI Speech to Text AI Subtitle Translation AI Text to Speech AI Video Summarizer Subtitle Remover Video Clipper
Pricing Documentation News About
Start Translating
Documentation

Get Started in Ten Minutes

Complete your first video translation task from scratch, plus detailed guides for all 7 AI tools

Quick Start

  1. Sign up / Log in: Click "Log in" in the top-right corner. New users get 60 free credits (enough for about 1 hour of video processing).
  2. Open Speech to Text or the Video Translation Workbench from the "AI Tools" menu.
  3. Drag in a video file, or paste a YouTube / TikTok / Bilibili link.
  4. Choose the video's source language (pick "Auto Detect" if unsure) and target language (34 supported).
  5. Choose a processing mode: Subtitles Only / Subtitles + Dubbing / Dubbing Only / Subtitle File Only.
  6. Pick a voice (we recommend "Auto"), then click Submit.
  7. Wait for processing to finish, then download your video / subtitles / dubbed audio.
Everything runs on cloud engines — no models to download, ready to use the moment you open the page.

7 AI Tools

Each tool is an independent page — submit and download separately, no interference between tasks:

ToolEntryWhat it does
Speech to Text/tools/transcribeUpload audio/video → transcribe speech → generate timestamped subtitles (SRT/TXT)
Subtitle Translation/tools/subtitleUpload SRT → one-click translate into 34 languages, timestamps preserved
Text to Speech/tools/ttsText → AI voiceover with 300+ voices, adjustable speed, MP3 output
Video Summarizer/tools/summarizeLong video/audio → auto-extracted key points and overview
Subtitle Remover/tools/videoBox-select to precisely mask hardcoded subtitles / watermarks
Video Clipper/tools/clipCut video segments by start/end timestamps
Voice Cloning/tools/cloneUpload a 10-second voice sample → clone a custom voice → saved permanently

Need full one-stop processing (transcribe → translate → dub → merge)? Use the "AI Video Processing Workbench" (/workspace/workbench).

️ Speech Recognition (ASR) Setup

The platform uses cloud recognition engines — industry-leading quality, no local models to download:

  • Alibaba Bailian Paraformer (default, most accurate): streaming paraformer-realtime-v2, 50+ languages, auto punctuation and number formatting. Requires a DashScope API Key.
  • Volcengine ASR: the same bigmodel tier used by Douyin. Requires AppID + Token.
  • Tencent Cloud ASR: recorded-file recognition. Requires SecretId + SecretKey.

Fill in the corresponding key in the "️ Cloud Service Settings" panel of the workbench. Every engine has a "Test Connection" button.

️ Engines without a configured key are unavailable — the task will prompt you to configure one rather than silently failing or falling back.

Translation Model Setup

Multiple LLMs supported — fill in the "️ Settings" panel of the workbench:

  • Kimi (kimi-k3): great at complex sentences and technical terms. Requires API Key.
  • DeepSeek: natural and cost-effective (recommended). Requires API Key.
  • OpenAI-compatible: works with Qwen / 302.AI / SiliconFlow etc. Requires Base URL + API Key + model name.
  • Baidu Translate: general API with free quota. Requires APP ID + secret.

The default "Auto" mode falls back in the order Kimi → DeepSeek → OpenAI → Baidu — if one fails, the next is used automatically.

Each model can be "Test Connected" individually to verify availability immediately after setup.

Dubbing Engine (TTS) Setup

  • Edge (default): Microsoft online neural voices, 300+ options, zero configuration.
  • MiniMax (cloud): top-tier Chinese dubbing, speech-02-hd with near-human emotion and quality. Requires MiniMax API Key + GroupId.
  • Alibaba Bailian CosyVoice (cloud): Tongyi Lab emotional voices, shares the DashScope Key with Bailian ASR.
  • OpenAI-compatible: any /v1/audio/speech endpoint (302.AI etc.). Requires Base URL + Key + model name.
  • Local TTS service: self-hosted service URL (needs GPU), data never leaves your network.

Configured cloud engines are automatically placed at the top of the voice dropdown; if Edge fails, it automatically degrades to a configured cloud engine — tasks never get stuck.

Each voice in the dropdown can be previewed; click Favorite to pin it to the top.

Voice Cloning

  1. Open the "Voice Cloning" tool page (/tools/clone).
  2. Upload a 10+ second clear voice sample (audio or video; audio is extracted automatically).
  3. Click "Clone" and wait (cloud MiniMax processing, about 1 minute).
  4. Cloned voices are saved permanently to your account and available in any dubbing task.
️ The cleaner and longer the sample, the better the clone. Use clean voice without background music or reverb.

Multi-Speaker Dubbing

Perfect for interviews, lectures, movie commentary and other multi-speaker videos.

  1. Turn Multi-Speaker Dubbing on in the dubbing settings.
  2. The system automatically identifies each speaker (male / female / child / elderly).
  3. Assign a voice to each speaker (leave blank for auto-assignment).
️ Speaker separation relies on a locally deployed pyannote model; if not deployed, it gracefully degrades to single-voice dubbing without affecting task completion.

Credits & Billing

Pay-as-you-go credits — no prepaid subscription, only pay for what you use:

  • 60 free credits on signup — enough to try the full workflow.
  • Credits are deducted by processing duration / character count — see "Account → Credit Ledger".
  • Low on credits? Buy a pack: 500 credits ($1.49) / 2000 credits ($4.99) / 8000 credits ($14.99).
Task TypeBilling BasisRate
Speech to Text / Subtitle Translation / Summarize / Subtitle Remover / ClipAudio/video duration1 credit / minute
Text to Speech / Subtitle DubbingCharacter count1 credit / 200 chars
Link download + subtitlesAudio/video duration2 credits / minute
Failed tasks are not charged; you can cancel anytime during processing.

FAQ

Why is the first task slow?

First use of a cloud engine establishes a connection and warms up; afterwards it returns to normal speed. Long videos also naturally take time for recognition and dubbing.

What's the maximum video length/size?

Browser uploads are limited to ≤ 4GB per file (recommended). Link downloads have no such limit.

Can I process in batches?

Yes. Enable " Batch Mode" in the workbench and select multiple videos — they queue and run automatically.

Is the translation quality industry-leading?

Recognition and translation use the same cloud engines as industry benchmarks (Alibaba Bailian / Volcengine / Tencent + Kimi / DeepSeek LLMs), with subtitle refinement support — output quality is fully comparable.

Is my data safe?

Recognition and dubbing use official vendor APIs with encrypted transmission. Only the necessary audio/video segments are uploaded for processing; nothing is stored or used for training.

What if I run out of credits?

Buy a credit pack in "Account → Recharge". Multiple payment methods supported; credits arrive instantly.