Google Supercharges Gemini Audio with 3.5 Transcribe, Adding Advanced Jargon and Speaker Separation
5
What is the Viqus Verdict?
We evaluate each news story based on its real impact versus its media hype to offer a clear and objective perspective.
AI Analysis:
The news presents measurable technical improvements (jargon handling, speaker separation) but applies them to an already advanced product line, making the overall impact moderate and highly targeted.
Article Summary
Google has updated its Gemini Audio capabilities with 3.5 Transcribe, marking a significant refinement of its voice processing suite. This new model automatically detects specialized jargon and improves transcription accuracy across over 85 languages. Key features include the ability to 'edit naturally with just your voice,' automatically removing filler words like 'um' and 'uh,' and allowing users to provide a customized vocabulary to handle unique industry jargon. Furthermore, it can attribute speech to up to three distinct speakers in pre-recorded audio and provide detailed word-level timestamps. These updates, which also include improved live modes (3.5 Live and 3.5 Live Experimental), improve reliability when faced with background noise or speech interruptions. The features are rolling out first for macOS and Android, with developer access available via the Gemini API.Key Points
- Gemini 3.5 Transcribe introduces sophisticated text cleaning, automatically removing filler words and structuring text post-transcription.
- The model enhances accuracy by allowing users to input custom vocabularies and specialized jargon, making it useful for technical or academic fields.
- It significantly improves multi-speaker handling, accurately attributing speech to up to three different voices and providing detailed timestamps.

