Any recorder
iPhone, Android, pro, old — format does not matter.
Dictaphone to text
Upload voice. Get short readable text.
iPhone, Android, pro, old — format does not matter.
Multi-hour lectures and meetings handled whole.
Interviews and dialogs labeled by who is speaking.
Text with punctuation and paragraphs — paste into Word or Notion.
Built for people who work from recordings
Turn meeting recordings into notes and follow-up text.
Review interviews with speakers, timestamps, and search.
Make transcript text for captions, posts, and show notes.
Convert lectures and webinars into study material.
In depth
A voice recorder saves the day when there's no time to type: a journalist holds it during an interview, a student sets it on the desk in a lecture, a researcher takes it into the field, a manager switches it on in a meeting. The recording comes out honest and complete, but the completeness is the problem. An hour of lecture is an hour you must replay to pull out a couple of definitions. A forty-minute interview turns into an evening of transcription under headphones with endless five-second rewinds. Field notes pile up as dozens of tracks, and finding the right moment means scanning them one by one. A recorder catches the moment perfectly, yet gives back no text — and without text you can't quote it in an article, drop it into a report or file it.
Upload the recorder's file by dragging it in — M4A from an iPhone, MP3 from Android, WAV from a pro recorder are all accepted as is, up to 4 GB per track. A multi-hour lecture needs no splitting: it's processed whole. The service recognises speech, adds punctuation, breaks it into paragraphs and, when several people speak, separates their lines by speaker — in an interview you immediately see who asked and who answered. Every fragment is tied to a timecode, so in the play-along editor it's easy to return to a doubtful term and fix it against the audio. The finished transcript exports to DOCX for handing in, to TXT for finding quotes, and on top of it a lecture summary or a list of meeting decisions is built.
The main enemy of a recorder track is the distance to the microphone and whatever covers it. Don't bury the device deep in a pocket or bag: fabric muffles the voice more than any compression. In an interview, place it closer to the person, not to you — you'll remember your own questions; the answers matter more. In a hall, sit nearer the speaker: a big room's echo blurs word endings. Set the main language before you start — for lectures full of terms and surnames it decides how much less you'll have to fix. And don't wipe the original right away: until you've checked the key figures and names against the timecodes, the source recording stays your insurance against a typo in an important place.
Pick a recording type, format, or workflow.
Transcription questions
Drag the recorder's file into the upload area — M4A, MP3 or WAV go in directly, with nothing to convert. The service recognises speech, adds timecodes and separates speakers if the recording is a dialogue. Then in the play-along editor you check terms and names against the audio and export a summary or an interview transcript.
Yes, a long track is processed whole, with no need to cut it into pieces. You get one continuous text with timecodes, and on top of it a summary with the key definitions and points. So you needn't replay the lecture: read the digest and open the exact moment by its timestamp.
The first recording can be transcribed free and without a card — handy for running a real interview or lecture and seeing how the text reads, how lines split and how cleanly your recorder is recognised. An account lets you store transcripts and gather them into knowledge bases by research topic.
Accuracy depends on conditions: a close mic, a quiet room and one speaker at a time give far cleaner text than a pocket recording in a noisy hall. There's no 100% guarantee, so check doubtful terms and surnames in the editor, replaying the original by its timecode.
Yes, diarisation marks the speakers' lines onto separate tracks, and the text shows the question-answer alternation. After processing you give the participants real names, and the whole interview transcript becomes a readable dialogue from which you can lift direct quotes with their timestamps.
The transcript exports to DOCX for handing in finished text, to TXT for searching the content and quotes, and to JSON for integrations. If the recording ran alongside video, SRT and VTT come in handy. One run produces every format at once, with nothing to rebuild by hand.
Vibe2Text
Upload audio or video while launch access is free. The start period has no duration cap.