廣東話 · Cantonese

Cantonese speech to text — transcribe 廣東話 audio

Speak Cantonese into the mic or upload Cantonese audio — WhatsApp voice notes, meetings, interviews — and get written Cantonese in the Chrome side panel.

Most speech tools treat Chinese as Mandarin and mangle Cantonese into the wrong words entirely. Whisper AI recognizes Cantonese as itself, in traditional characters — including written-Cantonese words like 嘅, 咗, 唔, and 佢 that Mandarin doesn't use. Hong Kong speech flips into English constantly — send email 俾我 — and the transcript keeps each language in its own script.

Two ways to use it

Record Cantonese speech, or upload Cantonese audio

Record your voice

Hit the mic and speak Cantonese naturally — the transcript comes out in traditional characters, ready to copy anywhere.

Upload an audio file

WhatsApp voice notes, recorded meetings, interviews — upload MP3, WAV, M4A, MP4, FLAC, or OGG.

How it works

Three steps, no setup

Open the side panel

Click the extension icon on any website. No language settings needed — Cantonese is detected automatically.

Speak or upload

Record live or drop in a file. Audio is processed securely and deleted immediately after transcription.

Copy your Cantonese text

Edit the transcript in the panel, then copy it into WhatsApp, Gmail, Docs — or download it as a file.

唔使再逐隻字打

No account needed. Works in 110+ languages besides Cantonese.

Add to Chrome

Cantonese transcription questions

Does it write Cantonese or convert to Mandarin?

Written Cantonese — 嘅唔咗佢 and friends, in traditional characters — not a Mandarin rewrite. That's the key difference from tools that only support 普通話.

Does Hong Kong-style English mixing work?

Yes — that's normal Hong Kong speech. English words stay in English; the Cantonese stays in characters.

Can I transcribe Cantonese WhatsApp voice notes?

Yes — download from WhatsApp Web (.ogg) and upload in the panel. Guide: Transcribe WhatsApp voice notes.

What about Mandarin?

Mandarin is fully supported too (simplified or traditional output) — auto-detect tells the two apart from pronunciation.