ElevenLabs Scribe V2
Transcribe long audio or video into structured text with language detection, timestamps, optional speaker labels, and sound-event tags. Choose Scribe
Transcribe long audio or video into structured text with language detection, timestamps, optional speaker labels, and sound-event tags. Choose Scribe V2 when downstream workflow steps need both the transcript and machine-readable word metadata. The response preserves text, detected language and confidence, duration, and optional timed speaker segments.
Media
StringAudio or video file to transcribe, up to the service limits.
Media
StringAudio or video file to transcribe, up to the service limits.
Language Code
StringISO-639 language code, or auto for language detection.
autoIdentify Speakers
BooleanLabel different speakers in the timed transcript.
falseMaximum Speakers
NumberExpected maximum speaker count, or 0 for automatic detection.
0Timestamp Granularity
StringWhether timing metadata is omitted, word-level, or character-level.
wordTag Audio Events
BooleanMark non-speech events such as laughter, footsteps, or applause.
trueClean Transcript
BooleanRemove filler words, false starts, and other disfluencies.
falseKey Terms
StringComma-separated names or specialist terms that should be favored.
Temperature
NumberSampling temperature, or -1 to use the model default.
-1Seed
NumberOptional random seed for repeatable transcription.
Response
ObjectStructured transcript with language and optional timing metadata.
Nodespell Team
Type
Node
Status
Official
Package
Nodespell AI
Category
AI / Audio / ElevenlabsInput
Output