
Real-Time Multimodal Interaction
Real-time voice and video experiences with streaming responses, stateful sessions, interruption handling, and seamless recovery.
- WebSockets
- Multimodal AI
- Streaming
- Session State
Staff AI Engineer · Google DeepMind
Staff AI Engineer | LLM Platforms & Backend Systems
Building production multimodal AI systems on the Gemini platform at Google DeepMind.
Real-time interaction · Agentic tool use · Grounded retrieval · Inference optimization · Distributed systems
San Diego, CA

Profile
Joon Shaw is a Staff AI Engineer focused on production-grade multimodal AI systems, LLM platforms, and backend infrastructure. Over more than 9 years at Google, he has worked across distributed backend systems, large-scale ML serving, and developer-facing AI platform capabilities. His current work at Google DeepMind focuses on turning frontier Gemini models into reliable services that external developers can build on, including real-time voice and video interaction, agentic tool use, inference optimization, grounded retrieval, and model evaluation.
He works across research, product, infrastructure, and developer-relations teams to define platform contracts, reliability standards, safety guardrails, and API surfaces for production AI workloads.
What I build
Production platform systems that help developers build with AI—from real-time multimodal interaction and tool-using agents to grounded retrieval, efficient inference, and secure, reliable APIs.

Real-time voice and video experiences with streaming responses, stateful sessions, interruption handling, and seamless recovery.

Agent orchestration systems that safely select tools, execute structured actions, coordinate parallel steps, and complete real-world tasks.

Retrieval systems that connect model responses to live search and trusted documents, producing answers backed by traceable evidence.

Model-serving infrastructure that improves latency, throughput, and operating efficiency through caching, batching, and workload-aware routing.

Secure, versioned developer APIs with authentication, usage controls, configurable safety, observability, and production reliability.
Career
More than nine years at Google, advancing from Software Engineer to Staff within Google DeepMind.
Oct 2022 – Present
Staff Software Engineer
Developer platform & runtime surfaces for production multimodal AI
Sep 2019 – Oct 2022
Software Engineer
Large-scale ML serving & distributed backend infrastructure
Aug 2016 – Sep 2019
Software Engineer
Backend services & distributed-systems foundations
Focus areas
Notes & insights
Technical notes on building reliable AI platform systems, real-time multimodal interfaces, agentic tool use, inference optimization, and grounded generation.
Coming soon
Notes on stateful streaming, interruption handling, speech detection, and session resilience for production voice AI systems.
Coming soon
Practical architecture patterns for function calling, schema constraints, parallel tools, and safe orchestration.
Coming soon
How context caching, serving tiers, batching, and workload-aware routing can reduce latency and cost.
Background
Bachelor of Science in Computer Science
2012 – 2016
Get in touch
For professional inquiries, technical collaboration, speaking, or writing opportunities, you can reach me by email.