Subtitling · translation · accessibility

Subtitle creation, transcription and timecoding

LinguaVox creates subtitle files from video or audio when a production does not yet have a reliable transcript or timed-text asset.

The service can include transcription, speaker identification, spotting, segmentation, timecoding, text editing, source-language template preparation and final file delivery.

We work with documentary, film, television, interviews, corporate communication, conferences, training, webinars and other audiovisual content.

Get a quote WhatsApp +34 637 822 394
Audiovisual editor working with waveform, video and subtitle segments
Since 2000Translation and audiovisual project management
languagesInternational language coverage
Quality managementDefined and traceable processes
Translation qualityIndependent bilingual revision when applicable
Human post-editingQualified review of machine translation

Client reviews

Starting from the finished video

A script is useful, but it is not essential.

When only the audiovisual file is available, we can work directly from the soundtrack and picture. The spoken content is identified and prepared for display as timed text.

This is not the same as producing a verbatim transcript. Speech contains repetition, hesitation, interruptions and incomplete structures that may not belong in a readable subtitle file.

The subtitle creator also has to decide where each event begins and ends, where sentences should be divided and how much text can reasonably be read while the viewer follows the image.

Transcription for subtitling

A subtitle transcript is prepared for a specific audiovisual purpose.

In conventional subtitles, the focus is generally on dialogue that the viewer needs to understand. In accessible timed text, relevant sound information and speaker identification may also need to be represented.

Audio quality has a direct effect on the work. A studio narration is easier to process than an outdoor interview with wind, several speakers and background noise.

Names, brands, technical terminology and place names should be checked against reliable reference material whenever possible.

Spotting and timecodes

Every subtitle event needs an in-time and an out-time.

Spotting links the text to the speech, shot structure and rhythm of the programme. Good timing helps the viewer read naturally without feeling that the subtitles are lagging behind or anticipating the audio.

Professional timed-text guidelines also recognise the importance of synchronising subtitles with both audio and image and of avoiding timing choices that disrupt shot changes or reveal dialogue too early.

The exact rules depend on the destination. A broadcaster, platform, festival or client may have its own timing specification.

Segmentation

Spoken sentences often have to be divided across several subtitle events.

Those breaks should follow the linguistic structure of the content wherever possible. Separating closely related words can make a subtitle noticeably harder to process even when the individual words are correct.

Segmentation also affects later translation. A well-prepared source template gives target-language translators a clear structure while still allowing them to adapt it to their own language.

Reading speed and text editing

The amount of text that can appear depends on the time available and the destination specification.

Professional guidelines typically define reading limits for particular workflows, but there is no useful universal number that should be imposed on every project regardless of platform, language or audience.

When speech is too dense for the available screen time, the subtitler may need to remove unnecessary repetition, reformulate a phrase or resegment the dialogue while preserving the information the viewer needs.

The aim is not to reproduce every sound mechanically. It is to create timed text that can actually be read.

Creating source templates for translation

For multilingual productions, subtitle creation is often the first stage of localisation.

We can prepare a reviewed source-language template containing the timed dialogue and relevant contextual information for later translation.

A template may include speaker names, annotations, terminology notes, on-screen text guidance and other information that helps downstream translators understand the scene.

It does not have to be a word-for-word transcript. A useful template should be accurate, readable and suitable as the basis for target-language subtitles.

Working from scripts

A script can speed up transcription and help verify names, but it should not automatically be treated as the final spoken text.

Actors, presenters and interviewees may change wording during recording. A documentary interview can depart substantially from any prepared questions or notes.

Where the service requires it, we compare the script with the final audiovisual material and use the recording as the reference for what is actually said.

Interviews and documentaries

Unscripted material presents specific challenges.

Speakers can change direction mid-sentence, correct themselves, overlap, use local references or switch between languages.

Documentaries may also contain archive extracts, captions identifying interviewees, historical terminology and multiple recording conditions.

Source creation has to make those elements clear enough for both the viewer and any translators who will later use the template.

Film and television

In scripted entertainment, timing interacts closely with editing and performance.

A subtitle that appears before a character speaks can reveal information too soon. A line held across a shot boundary may feel visually awkward. Fast exchanges may need careful event spacing so the viewer can follow who is speaking.

When a production has an established timed-text specification, we create the files to that specification rather than applying a generic house rule.

Corporate video and training

Business and training content often contains specialist terminology, product names and material displayed on screen at the same time as the dialogue.

The subtitle creator must consider this visual load. A long subtitle can obscure or compete with a software demonstration, diagram or presentation slide.

Client glossaries, scripts, product documentation and approved terminology can be used to check the source template before it is translated.

Multiple speakers and speaker identification

Some content makes it obvious who is speaking. Other material does not.

Off-screen voices, telephone conversations, panels, interviews and accessible subtitles can require speaker identification.

The format and wording of those identifiers depend on the project specification. A reliable participant or character list is useful reference material when available.

On-screen text

Dates, locations, signs, phone messages, interfaces and other on-screen information can be important to the viewer.

Not every visible word needs its own subtitle. The subtitler has to determine which information is relevant and how it should be handled when it appears at the same time as dialogue.

For multilingual projects, these elements can also be annotated in the source template so that translators know what requires localisation.

Speech recognition tools

Automatic speech recognition can be useful for generating an initial transcript, particularly with clear audio and predictable speech.

It can also make serious errors with names, technical vocabulary, accents, overlapping speakers and poor recordings.

Even a very accurate transcript is not a finished subtitle file. It still needs event timing, segmentation, reading control and audiovisual review.

We use automation where it provides a practical benefit and apply the human work required for the agreed deliverable.

Delivery formats

We can prepare common timed-text formats including SRT, WebVTT, STL, TTML and SCC.

The appropriate file depends on the editing system, publishing platform or distribution workflow.

For web video, WebVTT is specifically designed as a timed-text format associated with media and can carry subtitles or captions. Other professional environments use different formats and delivery specifications.

What we need to quote

Please provide the video or a representative sample where possible.

We also need the duration, source language, intended use, required format, target languages if translation will follow, and deadline.

Scripts, participant lists, glossaries and client specifications should be supplied at the start when available.

Frequently asked questions

Can you create subtitles when I only have the video?

Yes. We can create the timed subtitle file directly from video or audio. Reference materials are helpful but are not mandatory for every project.

Is transcription the same as subtitle creation?

No. A transcript records spoken content. Subtitle creation also requires timing, segmentation and adaptation for on-screen reading.

Can you create a master template for several languages?

Yes. This is a common way to organise multilingual projects. We can prepare a reviewed source or pivot template for subsequent translation.

Can you use my automatic transcript?

Yes, if it is suitable. We first assess and correct it as necessary, then create the timed subtitle events and complete the other work required for the final file.

Do you follow platform timing rules?

Yes, when the destination platform or client provides a specification. Different workflows can have different timing, line, reading and file requirements.

Can subtitle creation and translation be included in one quote?

Yes. We can quote the complete workflow from raw video through source creation and translation into one or several languages.

Send us your video, languages and deadline

We will review the material and confirm scope, delivery format, turnaround and price before work begins.

Get a quote WhatsApp