Echoon turns speech into polished, useful text — live from a microphone or meeting, or from a recording — with AI corrections, translation, summaries, and assistants on top. This guide covers what every mode shares first, then what makes each mode special.
You can look around anonymously, but transcribing requires an account: guest transcripts live only in the current browser session and guest accounts get neither free transcription nor free credits. Sign in — Microsoft (Azure AD work/school, recommended for companies), Google, Apple, GitHub, Facebook, X, or LINE — to transcribe free, receive your weekly free credits, keep your history, and sync across devices.
Credits are delivered the moment payment clears, so purchases are non-refundable as a rule — but a purchase whose credits you have not used at all is refunded in full within 14 days. Once any of them have been used, that purchase is no longer refundable; there is no partial refund. Only the purchased credits count: spending your weekly trial allowance or a reward after buying does not affect your refund, because those are used up first. The full rules are on the Refund Policy page, linked from the account menu.
Pick a mode from the Home screen or the sidebar, choose the recognition language, and press start (or drop a file). Every session is saved to History in the sidebar, where you can search, rename, share, or delete it.
The setup panel in front of each live mode (and the settings drawer during a session) is built from the same blocks. A minute spent here pays off directly in accuracy.
Echoon recognizes 29 languages. Set your default in Settings → Speech recognition; each session can override it, and Bilingual Talk detects the language of each sentence automatically.
Live modes can record your microphone, the speaker/system audio, or both — so an online call is captured from both sides, not just yours. Speaker capture opens a screen-share prompt: pick the tab or screen playing the audio and tick "Share audio".
Tell the AI where the conversation happens (one scene) and which industries it touches (multi-select — an IT project for a hospital can be both). Curated term lists load, and corrections, translation wording, and note-taking all adapt. Both default to Auto: leave them there and the AI names the setting from the conversation itself once it has heard enough, then works from it for the rest of the session. Add your own scenes and industries in Settings; any language works.
Names, products, and jargon come out spelled right when recognition is biased toward them. Use the per-session hotwords box for one-off terms, and Settings → Dictionary for the terms you always need.
Paste context or upload a PDF, Word, or text file — an agenda, resume, or product doc. Text is extracted on your device, long documents are AI-compressed into a digest, and everything grounds the AI: corrections, Q&A, insights, and speaking suggestions all get sharper.
Turn on speaker labels to separate who said what. Enroll a voice once in Settings → Speech recognition (read the sample passage for 10–30 seconds) and transcripts show the real name instead of "Speaker 1" — you can also enroll a speaker straight from a finished transcript. Extra samples per person make matching noticeably more reliable.
In Advanced and Meeting sessions the assistant works in real time. Standard, Voice Memo, and File transcripts get the same panel after a one-tap AI enhancement.
A structured key-point digest — decisions, action items, open questions, and more, with sections that adapt to your scene — builds while people are still talking and is finalized the moment the session stops.
Ask anything about what has been said so far. With question auto-detection on, Echoon notices when the other side asks you something and drafts an answer automatically; every exchange is saved to the Q&A tab.
Stuck for words mid-conversation? One tap gets reply suggestions grounded in your background documents and the conversation so far.
One tap analyzes the conversation to date — open questions, risks, and action items, tailored to the scene and your background material.
Any finished Standard, Voice Memo, or File transcript can be upgraded afterwards with sentence corrections, translation, and a summary. Flip on auto-enhance in the setup panel to run it automatically when the session ends. If enhancement fails, nothing is charged.
Six ways in, all built on the same foundation — pick by how the audio arrives and how much AI you want on top.
Plain realtime transcription of your microphone and/or speaker audio — the fastest, cheapest way to get accurate words on the page.
Realtime transcription plus the full AI copilot: sentence-by-sentence corrections, live translation, a running key-point digest, and Q&A, insights, and speaking suggestions on demand.
A realtime two-way conversation between two languages — put the screen between you and whoever you do not share a language with.
Paste a Zoom, Google Meet, or Teams link and an AI notetaker bot joins the call, relaying the meeting audio into the Advanced pipeline — corrections, translation, digest, and all assistants included.
Record locally in the browser, then transcribe in one batch pass when you stop.
Upload an existing audio file and get the full transcript back at the Standard rate.