Mistral Worldwide Hackathon, Online Edition

Yongkang Zou · 2026-03-02 · hackathon · Emotion & Vision AI

AI can now transcribe my awkwardness, not just my broken English.

As AI reasoning keeps improving, we can safely say model IQ is advancing at an incredible speed. But when it comes to human-computer interaction, building better EQ is a different problem. Multimodal intelligence is becoming a real necessity for this kind of interaction.

That is why we spent this weekend at the Mistral Worldwide Hackathon building Evoxtral. We built it on top of Voxtral-Mini-3B-2507. Using over 1000 expressive speech samples for SFT and RL training, we created a model that reads your speech and reads between the lines in real time.

Understanding the "how"

Imagine an AI interview system that evaluates not just your technical answers, but your hesitation, confidence, and behavioral cues. This same capability unlocks richer meeting intelligence, sharper sales call reviews, and voice agents with far better EQ. We need systems that understand how you say something, not just what you say.

I had huge fun building this over the weekend with my friends Aetos Huo and Lior Li.

You can try our model on HuggingFace:
mistral-hackaton-2026/evoxtral

Or play with the live demo in our studio app:
Evoxtral Live Demo

Huge thanks to Iterate and Mistral AI for organizing an amazing hackathon and making it a truly worldwide event. Thanks also to Weights & Biases, Hugging Face, and ElevenLabs for the sponsorship. Shoutout to Simon Lorenzo and the team for making it all happen.

Testing the Evoxtral live demo

Evoxtral analyzing speech patterns in real time