{Tech:Europe} Paris Hackathon: Rewind

Yongkang Zou · 2026-02-22 · hackathon · Spatial Intelligence

This was my last hackathon in Paris before moving to Stockholm. I built Rewind with Alessandro Lombardo and Christian Wallenwein at the {Tech: Europe} Paris AI Hackathon.

Rewind turns a single photograph into an explorable 3D world with ambient sound and a voice companion that helps you reflect on the memory. You upload a photo, and the system reconstructs the scene as a navigable environment you can walk through with keyboard controls, while an AI guide asks you about the moment.

It's a toy project, but the idea came from something real. Leaving Paris meant seeing Alessandro and Christian less often. We wanted to build something about preserving moments, about being able to step back into a place and a feeling. Rewind is that: a way to keep and relive the memories that matter.

How It Works

Five steps from photo to explorable memory:

Child riding a bike
Memory input example
Child on a playground toy
Another memory input
  1. Upload a photo with context. You describe when and where it was taken, who was there.
  2. Scene reconstruction. The photo is converted into a navigable 3D world using LingBot-World.
  3. Free exploration. Walk through the reconstructed environment with WASD keyboard controls.
  4. Ambient audio. AI-generated soundscapes match the setting (waves for a coastal scene, wind for mountains).
  5. Voice dialogue. An AI companion guides reflection on the memory through conversation.

The Stack

graph LR
Photo[Upload Photo] --> Analyze[GPT-4o-mini: Scene Analysis]
Analyze --> LB[LingBot-World: 3D Generation]
Analyze --> Audio[ElevenLabs: Ambient Sound]
LB --> Scene[Explorable 3D World]
Audio --> Scene
Scene --> Voice[Gradium STT/TTS + GPT-4o-mini: Voice Companion]
  • Frontend: React + TypeScript + Vite + Tailwind CSS
  • Backend: FastAPI (Python)
  • 3D Generation: LingBot-World (Wan2.2) on Modal, 4x A100-80GB GPUs
  • Image Processing: fal.ai nano-banana for aspect ratio conversion
  • Scene Analysis: OpenAI GPT-4o-mini
  • Audio: ElevenLabs Sound Generation API
  • Voice Interface: Gradium STT/TTS with GPT-4o-mini for dialogue

LingBot-World: The World Model

The most interesting technical piece is LingBot-World, an open-source world model by Robbyant (Ant Group). It starts from Wan2.2, a 14B parameter image-to-video diffusion transformer, and extends it into a mixture of experts DiT with ~28B total parameters (14B per expert, one active per denoising step).

What makes it work for Rewind:

  • Action-conditioned generation. Keyboard and mouse inputs drive the evolution of future frames, so walking through a scene feels responsive.
  • ~16 FPS throughput with sub-second interaction latency.
  • Long-term consistency: up to 961 frames (one minute at 16 FPS) while keeping characters recognizable and environments anchored.
  • Trained on hybrid data: curated web videos + Unreal Engine synthetic data for diverse real-world scenes.

We ran it on Modal with 4x A100-80GB GPUs. The generation takes about 30 to 60 seconds to produce an explorable scene from a single photo.

HunyuanVideo for the Atmospheric Layer

We experimented with Tencent's HunyuanVideo (13B parameters, the largest open-source video generation model) for generating ambient video textures: sky movement, water rippling, foliage swaying. The 1.5 version with 8.3B parameters runs on consumer GPUs and generates in 8 to 12 steps (75% faster than v1 on RTX 4090).

In the final build, LingBot-World handled both the 3D navigation and the atmospheric effects, but HunyuanVideo was our backup for generating the "living" environment textures.

The Voice Companion

The part that ties Rewind together emotionally isn't the 3D. It's the voice.

As you walk through the reconstructed scene, an AI companion speaks to you. It knows the context you provided (who was there, when it was, why it mattered) and gently prompts reflection: "What were you feeling in this moment?" "Who was standing next to you?" "What would you tell them now?"

The pipeline: Gradium handles speech-to-text and text-to-speech, and GPT-4o-mini generates the conversational responses. Low latency was critical. Nobody wants to wait 5 seconds for a reflective prompt. The STT/TTS round-trip stays under 2 seconds.

Building in 24 Hours

Alessandro handled the 3D pipeline and Modal GPU orchestration. Christian built the frontend experience and audio integration. I worked on the backend, scene analysis, and voice companion.

The hardest part was getting LingBot-World to produce consistent, explorable scenes from arbitrary photos. A selfie at a cafe gives very different depth information than a wide outdoor shot. We used GPT-4o-mini as a pre-processing step. It analyzes the photo, describes the scene layout, estimates depth, and provides structured context that helps the world model produce more coherent environments.

A Different Kind of Souvenir

Rewind isn't going to production. It's a hackathon project, a proof of concept.

But building it was the point. My last weekend hacking with friends in Paris before moving to a new city. The irony wasn't lost on us. We built a tool for preserving memories while creating one.

Alessandro and Christian are still in Paris. I'm in Stockholm now. We talk less, meet less. That's what happens. But I have the photos, and now I know it's technically possible to walk back into them.

Source code: github.com/inin-zou/Rewind


Built at the {Tech: Europe} Paris AI Hackathon with Alessandro Lombardo and Christian Wallenwein. February 2026.

References

Rewind demo exploration