The starting point for any good interview is a full transcript of the conversation with a cabinet minister, a film star, the league’s top scorer or a prize-winning author. Once every word is on the page, it becomes far easier to produce a clear, accurate edit. More importantly, that edit will reflect the interviewee’s actual language and line of argument.
The problem is time. Turning audio into text is slow, often painfully so, and increasingly out of step with the demands of modern journalism, especially online.
Here, artificial intelligence offers a way through. One example is the Dutch start-up Amberscript B.V., which uses AI to automate transcription. Reporters upload their recording to the Amberscript website or app and receive a verbatim transcript within seconds. The system can also distinguish between speakers in a conversation – for instance, the journalist and the interviewee – and it works in 39 languages, including Italian.
Once the transcript is ready, the workflow speeds up again. By placing the cursor at any point in the text, a second cursor automatically jumps to the corresponding point in the audio file. The platform accepts audio files in M4A, MP3 and WAV, as well as video files in M4V, MOV and MP4. Transcripts can be exported in a range of formats, including Word, JSON, SRT, VTT, EBU-STL and plain text.
Amberscript’s services can be tested free of charge. After the trial, transcribing up to one hour from an audio or – crucially – a video file costs €15. A subscription of €40 a month covers up to five hours of transcription. The company also offers video subtitling, although at significantly higher rates: from €10 to €25 per minute, depending on whether the subtitles remain in the original language or are translated into another.