ClipCanva

Gemini 3.5 Transcribe Workflow: Verbatim vs Smart Transcripts, Then Scripts and Shorts

Google’s Gemini 3.5 Transcribe (26 Aug 2026) splits live captions from file transcripts. Use verbatim vs smart modes, then rewrite scripts on ClipCanva.

August 29, 2026ClipCanva Editorial

Gemini 3.5 Transcribe Workflow: Verbatim vs Smart Transcripts, Then Scripts and Shorts

A Gemini 3.5 Transcribe workflow is not “dump the MP3 into a chatbot and hope the Short writes itself.” Google’s 26 August 2026 product post ships a dedicated speech-to-text model with two desks: gemini-3.5-transcribe-live for sub-second streaming, and gemini-3.5-transcribe for pre-recorded files with speaker labels and word timestamps. The model card (updated the same day) and the audio transcription guide pin the operator limits: 85+ languages, smart vs verbatim modes, custom vocabulary, and a hard rule that smart cleanup cannot run with timestamps or diarization. ClipCanva is not affiliated with Google, Canva, Descript, or VEED, and Gemini 3.5 Transcribe is not a live ClipCanva generator. Transcribe on Google’s API or AI Studio, then bring the cleaned spine into ClipCanva’s AI video summarizer and AI script generator.

Official facts for Gemini 3.5 Transcribe and nearby tools

Source Public claim Workflow implication
Introducing Gemini 3.5 Transcribe (26 Aug 2026) Dedicated speech-to-text. Live API: gemini-3.5-transcribe-live. File API: gemini-3.5-transcribe. Artificial Analysis WER 4.0% streaming / 2.6% non-streaming. Time to final transcription 70% faster than Chirp 3. FLEURS: 5.50% streaming / 5.04% non-streaming WER. Smart transcription handles self-corrections, strips filler, auto-formats. Custom vocabulary. 85+ languages. Public preview in Gemini API / AI Studio; Rambler on Android; Gemini app on macOS (English). Chrome “coming soon.” Pick live vs file before you record. Do not treat Rambler dictation as a podcast pipeline.
Gemini 3.5 Transcribe model card Unary audio up to 1 hour; 30 minutes when diarization or word timestamps are on. Live session 10 minutes. Live: language auto-detect and smart formatting; no word timestamps, no speaker diarization. File: timestamps and diarization supported. Custom vocab up to 1,000 terms (best results typically up to 100). File diarization up to 8 speakers; 3+ experimental. Function calling, thinking, image generation, caching, batch, and flex/priority inference are not supported on this model. A 45-minute interview with speaker labels is a file job under the 30-minute feature cap, not a live stream.
Audio transcription guide Default mode is verbatim. Smart mode cannot be combined with timestamp_granularities or diarization_mode. Word timestamps may degrade accuracy. Inverse text normalization examples include spoken money becoming $26M. Run two passes when you need both a readable script and a timed caption file.
Gemini API pricing (updated 28 Aug 2026) Paid gemini-3.5-transcribe: audio input $2.00 / 1M tokens (~$0.003/min); text output $12.00 / 1M (~$0.002/min); blended ~$0.005/min. Paid gemini-3.5-transcribe-live: input $3.50 / 1M (~$0.005/min); output $21.00 / 1M (~$0.004/min); blended ~$0.009/min. Free tier exists; paid-tier audio is not used to improve Google products. Price the minute and the mode. Live captions cost more than a file dump.
Descript Upload or record, get a transcript, then edit the video by editing the text. Clips, captions, and translation sit in the same editor. A transcript that never leaves the editor is a cut, not a script brief.
VEED auto subtitles Auto-generate captions, restyle them, burn them in, or download SRT/VTT/TXT. Positions captions as essential because most mobile video is watched without sound. Subtitles are a delivery layer. They are not a YouTube script.

These are vendor-documented limits, not a promise that your SKU name, guest’s legal name, or quoted claim will survive smart cleanup. Access is public preview. Region, safety filters, and rate limits still sit on the live Google project.

Pick the job before you upload audio

Gemini 3.5 Transcribe fails when every brief is treated as one “transcribe this” box. The 26 August split is four different desks.

Job Where it actually runs What you must supply What the model may invent Fail condition
Live captions / voice agent gemini-3.5-transcribe-live, Live API, ≤10 min session Clean mic, language hint if known Partial lines with sub-second latency Asking for speaker IDs or word timestamps on the live socket
Verbatim archive gemini-3.5-transcribe, mode verbatim File ≤1 hour (≤30 min if you also want speakers or word times) Exact fillers, false starts, “um” Shipping the raw dump as show notes
Speaker-labeled interview File API, diarization_mode: "speaker" Distinct mics if you can; custom vocab for names spk_1 / spk_2 labels (3+ speakers experimental; card allows up to 8) A four-person panel treated as production-ready without a human pass
Timed captions File API, timestamp_granularities: ["word"] One language hint, no smart mode Start/end offsets per word Combining smart cleanup with timestamps (docs block it)
Readable script spine File API, mode "smart" Custom vocab for brand and SKU Filler gone, self-corrections resolved, lists and money formatted Using smart output as a legal transcript
ClipCanva rewrite AI video summarizer, AI script generator, YouTube script generator The cleaned spine plus the job (Short, explainer, podcast recap) A new hook, CTA, and scene list Pasting gemini-3.5-transcribe into a ClipCanva box
Burned-in social captions VEED or similar subtitle tools The approved SRT, not the smart essay Styled on-screen type Treating 99.9% marketing copy as a measured WER

If the brief is “this 28-minute product webinar must become three Shorts,” run verbatim + speakers on a file under 30 minutes, then a second smart pass only on the sections you will rewrite. If the brief is “live captions in a 8-minute AMA,” stay on the Live API and do not ask for spk_1.

The 26 August blog described attribution for up to three speakers, with 3+ experimental. The model card is stricter on the number and looser on the cap: file diarization supports up to eight speakers, and attribution for three or more is experimental. Quote the card in a client estimate, then verify the live response on a 60-second sample before you price a panel.

Delivery card before you pay per minute

Write the card before you hit transcribe. Official blended file pricing is about $0.005 per minute. A 30-minute webinar is about $0.15 of transcription before you spend anything on a rewrite. Live blended rate is about $0.009 per minute.

Job: [live captions | verbatim archive | speaker interview | word-timed captions | smart script spine]
Surface: [AI Studio | Gemini API gemini-3.5-transcribe | gemini-3.5-transcribe-live | Rambler | macOS Gemini app]
File length: [mm:ss]. If speakers or word times: must be ≤30:00.
Mode: verbatim OR smart. Never both in one call.
Language: auto, or BCP-47 hint [en-US / ja-JP / …]
Custom vocab (≤100 terms for best results): [SKU, guest name, product line]
Success: names spelled right, money/IDs intact, speakers labeled if promised, no invented claims.
Fail if: smart mode ate a quoted price, timestamps missing on a caption job, or a fourth speaker is trusted without review.

Starter custom vocabulary for a creator webinar (original; not copied from Google samples):

ClipCanva, Veo 3.1, Seedance 2.5, North Harbor 01, SKU-NH-30, BigQuery

Do not stuff everyday words. The transcription guide is explicit: bias domain terms, acronyms, and proper nouns. A 1,000-term dump is allowed and usually worse than 20 names you actually say.

Two passes beat one “do everything” call

Smart mode cannot carry timestamps or speaker labels. Word timestamps can hurt accuracy. That is not a bug in your prompt. It is the documented compatibility matrix.

Work it as a stack:

  1. Archive pass. verbatim on the full file. Add speaker labels if it is a conversation and the file is ≤30 minutes. Save this as the legal-ish source. Keep fillers if you need to prove what was said.
  2. Caption pass. Only if you need SRT. Same file, verbatim plus word timestamps, no smart mode. Expect a slightly noisier transcript. Fix brand spellings by hand.
  3. Spine pass. Smart mode on the same audio, or on the verbatim text in a separate writing model. This is the readable version: self-corrections resolved, “ums” gone, money formatted.
  4. Script pass. Move the spine into ClipCanva. Summarize the argument on the AI video summarizer. Rewrite the hook, proof, and CTA on the AI script generator or YouTube script generator. If the source is an episode, plan clips on the AI podcast generator.
  5. Picture pass. Only after the line is locked. Pull shot language from prompt ideas, then test a visual on the AI video generator or image-to-video.

Descript’s pitch is the opposite order: transcribe so you can cut the video by deleting words. VEED’s pitch is captions that survive mute autoplay. Keep those jobs on those tools. Gemini 3.5 Transcribe is the speech-to-text layer. ClipCanva is the script and generation layer. Mixing them in one prompt produces a fluent paragraph with the wrong SKU.

What smart mode is allowed to change

Google’s guide gives the contract. Verbatim keeps “um,” “uh,” “like,” repetitions, and false starts. Smart strips those, resolves “Tuesday—no, Wednesday,” and can turn spoken structure into paragraphs, lists, dates, and currency.

Use smart when you are writing:

  • show notes
  • a Short hook
  • an explainer voiceover
  • a recap email

Stay on verbatim when you are:

  • quoting a customer in a case study
  • logging a price, SKU, or medical claim
  • building captions that must match the mouth
  • handing legal or compliance a record of what was said

If a guest says “twenty-nine dollars, wait, thirty-nine,” smart mode may output $39. That is useful for a caption. It is dangerous if your landing page still sells at $29. Read the money lines against the waveform, not against the pretty paragraph.

Function calling is a different product surface. The 26 August post says the Gemini macOS app can delegate image generation and file analysis by voice. The transcribe model card lists function calling as not supported. Do not brief a client that gemini-3.5-transcribe will “also generate the thumbnail.”

Creator / operator checklist

  • [ ] Model id is gemini-3.5-transcribe or gemini-3.5-transcribe-live, not a general Gemini chat model asked to “transcribe this.”
  • [ ] Live jobs are ≤10 minutes and do not request speaker IDs or word timestamps.
  • [ ] File jobs with speakers or word times are ≤30 minutes; plain unary files stay ≤1 hour.
  • [ ] One mode per call: verbatim or smart. No timestamps and no diarization on the smart call.
  • [ ] Custom vocab is a short list of names and SKUs, not a paragraph.
  • [ ] Language hint is set when you already know the locale.
  • [ ] Quoted prices, legal names, and CTAs were checked against the verbatim pass.
  • [ ] 3+ speaker labels are treated as experimental.
  • [ ] Caption SRT came from the timed verbatim pass, not from smart prose.
  • [ ] Script rewrite ran on ClipCanva’s summarizer / script generator, not inside the STT call.
  • [ ] Visual tests wait until the spoken line is locked.
  • [ ] Client deck does not call Gemini 3.5 Transcribe a ClipCanva model.

FAQ

Is Gemini 3.5 Transcribe a video generator? No. It converts speech to text. Video generation is a different Google family (Veo / Omni). On ClipCanva, generation still goes through the AI video generator or image-to-video after the script exists.

Can I get smart cleanup, speaker labels, and word timestamps in one request? Not according to the transcription guide. Smart mode cannot combine with timestamp_granularities or diarization_mode. Run an archive pass and a spine pass.

How many speakers are supported? The model card: file diarization up to eight speakers; attribution for three or more is experimental. Live streaming does not support diarization. The product post described up to three speakers with 3+ experimental. Sample a real panel before you promise labels in a SOW.

What does it cost? Google’s paid-tier table (28 August 2026) lists a blended ~$0.005 per minute for gemini-3.5-transcribe and ~$0.009 per minute for live. Free-tier audio may be used to improve Google products; paid-tier audio is not. Re-read the live pricing page before you quote a volume job.

Is this on ClipCanva? No. ClipCanva is independent. Use Google AI Studio or the Gemini API for the speech-to-text step, then move the cleaned spine into AI video summarizer, AI script generator, or the YouTube script generator.

Google will keep moving the model card. Re-read the live transcription docs before you quote duration, speaker count, or price. The 26 August 3.5 Transcribe controls are real; they are still not a caption SLA on your guest’s name.