Joon Shaw

Staff AI Engineer · Google DeepMind

Joon Shaw

Staff AI Engineer | LLM Platforms & Backend Systems

Building production multimodal AI systems on the Gemini platform at Google DeepMind.

Real-time interaction · Agentic tool use · Grounded retrieval · Inference optimization · Distributed systems

San Diego, CA

Joon Shaw working at his desk, reviewing an AI architecture diagram and code on a monitor
Production AI platform engineering
  • 9+ years at Google
  • Staff AI Engineer
  • Gemini platform systems
  • Multimodal AI infrastructure

Profile

About

Joon Shaw is a Staff AI Engineer focused on production-grade multimodal AI systems, LLM platforms, and backend infrastructure. Over more than 9 years at Google, he has worked across distributed backend systems, large-scale ML serving, and developer-facing AI platform capabilities. His current work at Google DeepMind focuses on turning frontier Gemini models into reliable services that external developers can build on, including real-time voice and video interaction, agentic tool use, inference optimization, grounded retrieval, and model evaluation.

He works across research, product, infrastructure, and developer-relations teams to define platform contracts, reliability standards, safety guardrails, and API surfaces for production AI workloads.

What I build

Selected Platform Contributions

Production platform systems that help developers build with AI—from real-time multimodal interaction and tool-using agents to grounded retrieval, efficient inference, and secure, reliable APIs.

Voice and video streams entering an AI system that produces responses and adapts when the user interrupts.

Real-Time Multimodal Interaction

Real-time voice and video experiences with streaming responses, stateful sessions, interruption handling, and seamless recovery.

  • WebSockets
  • Multimodal AI
  • Streaming
  • Session State
Calendar, search, and email tools connected to an AI orchestration system that produces a verified result.

Agentic Tool Use & Function Calling

Agent orchestration systems that safely select tools, execute structured actions, coordinate parallel steps, and complete real-world tasks.

  • AI Agents
  • Function Calling
  • MCP
  • Structured Output
A question and search sources entering an AI retrieval system connected to documents and a verified answer.

Grounded Retrieval & Answering

Retrieval systems that connect model responses to live search and trusted documents, producing answers backed by traceable evidence.

  • RAG
  • Search Grounding
  • Reranking
  • Citations
AI requests routed through caching, batching, and priority-serving paths before reaching a model-serving system.

Inference Efficiency & Model Serving

Model-serving infrastructure that improves latency, throughput, and operating efficiency through caching, batching, and workload-aware routing.

  • Model Serving
  • Context Caching
  • Batch Inference
  • Latency
An application request passing through security, versioning, rate-control, safety, and monitoring checks before receiving a validated response.

Secure AI Platform APIs

Secure, versioned developer APIs with authentication, usage controls, configurable safety, observability, and production reliability.

  • API Design
  • Authentication
  • Safety
  • Reliability

Career

Experience

More than nine years at Google, advancing from Software Engineer to Staff within Google DeepMind.

  1. Google DeepMind

    Now

    Oct 2022 – Present

    Staff Software Engineer

    Developer platform & runtime surfaces for production multimodal AI

  2. Google

    Sep 2019 – Oct 2022

    Software Engineer

    Large-scale ML serving & distributed backend infrastructure

  3. Google

    Aug 2016 – Sep 2019

    Software Engineer

    Backend services & distributed-systems foundations

Focus areas

Expertise

Generative AI & LLM Platforms

  • Gemini API
  • Google AI Studio
  • Multimodal AI
  • LLM platforms
  • Real-time AI
  • Model evaluation
  • Grounded generation

Agentic Systems

  • Function calling
  • Tool-use orchestration
  • MCP support
  • Schema-constrained decoding
  • Parallel and chained tool calls
  • Agent architecture

ML Serving & Inference Optimization

  • Model serving
  • Context caching
  • Batch serving
  • Priority serving
  • Flex serving
  • Accelerator-aware routing
  • Cost and latency optimization

Distributed Systems & Backend Infrastructure

  • WebSockets
  • gRPC
  • REST APIs
  • Multi-region systems
  • Event-driven pipelines
  • Progressive rollouts
  • Observability
  • SLOs

Safety, Governance & Platform Quality

  • API contracts
  • Versioning
  • Deprecation policy
  • Scoped keys
  • Ephemeral tokens
  • Safety guardrails
  • Groundedness evaluation

Notes & insights

Writing

Technical notes on building reliable AI platform systems, real-time multimodal interfaces, agentic tool use, inference optimization, and grounded generation.

Coming soon

Building Reliable Real-time Voice AI on Gemini

Notes on stateful streaming, interruption handling, speech detection, and session resilience for production voice AI systems.

Coming soon

Designing Agentic Tool Use at Production Scale

Practical architecture patterns for function calling, schema constraints, parallel tools, and safe orchestration.

Coming soon

Inference Optimization Patterns for Large Multimodal Models

How context caching, serving tiers, batching, and workload-aware routing can reduce latency and cost.

Background

Education

Brown University

Bachelor of Science in Computer Science

2012 – 2016

Get in touch

Contact

For professional inquiries, technical collaboration, speaking, or writing opportunities, you can reach me by email.