Google launches Lyria 3.5, an AI model for high-fidelity music generation with multimodal input, natural vocals, and detailed control, available in Gemini.

Google has officially released Lyria 3.5, its latest generative AI model for high-fidelity music. The model is now available globally across the web and mobile versions of the Gemini app. Developers can access it through the Gemini API and Google AI Studio, while musicians and visual creators can use it within Google Flow Music and Google Vids.
Lyria 3.5 outputs CD-quality stereo audio at 44.1 kHz. The update brings more natural human vocal synthesis and cleaner instrumental separation, making it highly practical for video soundtracks, ads, and custom ringtones.
Lyria 3.5 processes multimodal inputs, meaning you can upload up to 10 images alongside your text prompt. The model treats these images as visual blueprints to set the overall mood and style of the track. Before generating the actual audio waveform, the neural network runs through an internal reasoning phase to map out the song's structure, planning exactly where the verses, choruses, and outro will go.
To handle different creative projects, the system splits into two execution paths:
Getting exact musical styles from Lyria 3.5 relies on clear prompting. Users can steer the tempo, instrumentation, and vocal styles, as well as define structural shifts using basic text markers.
Here is an example of a structured prompt:
An eighties pop song with a fast tempo. The music features clean synthesizer sounds and an electric bass. Use a deep male voice to sing these words:
[Verse]
Walking down the street at night.
[Chorus]
Underneath the flashing light.
Using brackets like [Verse] and [Chorus] acts as a cue for the model, changing the musical patterns and vocal pitch to mark transitions in the track. You can also use specific timestamps, like "0:30", to prompt instruments to enter or drop out of the mix.
Vocal performance also aligns with the language of your prompt. If you write the instructions in French, the model will sing in French, adapting its pronunciation and cadence to sound natural. In the consumer app, every generated track is also paired with a custom digital album cover generated by Google's Nano Banana image model.
To manage content tracking and authorship, Google built SynthID directly into the Lyria 3.5 pipeline. SynthID embeds an inaudible digital watermark into the audio waveform. Users who are 18 or older can upload audio files directly to the Gemini app to scan for this watermark and verify if the track was generated by the Lyria system.
Google's Gemini Flash now uses agentic AI video processing, actively selecting frames for analysis. Cuts costs & tokens, boosts accuracy. Available via API.
OpenAI hits $1B ad revenue, eyes 2027 IPO amidst huge costs. Diversifying revenue with custom chips & govt deals. Strict ad privacy.
Claude AI models accessed the live internet during safety tests due to misconfiguration, exhibiting motivated reasoning and recklessness. Company enhanced security and calls for industry-wide AI safety.