Skip to content
AI•9 min read

Antigravity SDK with Local Models: Your GPU, Your Tokens, Zero Cloud

At 12:14 AM on a Sunday, my content pipeline ran out of quota. It wasn't a bug. The API returned a 429 halfway through the batch and the job died with 18 posts half-generated. I sat there staring at the log with cold coffee, realizing I had just mortgaged my night to a rate limit I don't control.

A month later, the feature that would have saved that night landed. The Antigravity SDK now runs agents on a local model. No cloud, no API key, no rate limits.

"Today, we're announcing that the Antigravity SDK supports local workflows across a wide range of local models and execution options, featuring initial support for Gemma 4 26B A4B using Google AI Edge's LiteRT." - Source: Google Developers Blog

This post walks through the install, your first local agent, and the traps the announcement doesn't put in bold. With real numbers and a clear rule for when local wins and when the cloud does. I'm not here to convince you your laptop is a datacenter.

Image: Google Developers Blog (original article)

What Google actually announced

The announcement is dated September 23, 2026, and it's signed by Sachin Kotwani and Taylor Mullen. Three pieces fit together here:

  • The Antigravity SDK: the Python library (pip install google-antigravity) for building agents with the same agentic capabilities Google Antigravity uses.
  • LiteRT-LM: Google AI Edge's orchestration layer for running LLMs on-device, cross-platform, with GPU/NPU acceleration.
  • Gemma 4 26B A4B: the model Google tested all of this with, packaged as a .litertlm file.

The result: you can give your code agentic assistance completely offline.

"With this new support you can enable agentic assistance via local models completely offline." - Source: Google Developers Blog

One detail from the SDK README changes your expectations. The PyPI package ships a compiled runtime binary. Cloning the repo won't run it.

"The Google Antigravity SDK relies on a compiled runtime binary that is included in the platform-specific wheels published to PyPI. Cloning this repository alone is not sufficient to run the SDK." - Source: antigravity-sdk-python

Why run the agent on your machine

Google sums it up in four reasons, and none of them are empty marketing:

  • Cost: no per-token bills and no rate limits.
  • Privacy: your code never leaves your machine, useful under strict compliance requirements.
  • Offline resiliency: the agent works without a stable internet connection.
  • Hybrid workflows: the cheap, repetitive work locally; the heavy lifting in the cloud when needed.

"Cost efficiency: Execute local agentic workflows without API costs or rate limits." - Source: Google Developers Blog

Privacy is the one that matters most to me on client work. When you're auditing billing.py for a bank, uploading the file to an API isn't a legal option, and signing a DPA doesn't undo the fact that the code traveled. A local agent fixes that at the root: the file doesn't move.

Install and your first local agent

Google recommends a machine with more than 24 GB of VRAM or unified memory. With that in mind, the install is three commands.

# terminal
python3 -m venv .venv
source .venv/bin/activate

pip install google-antigravity litert-lm

# Fetch the ~26B model from Hugging Face into ~/.litert-lm/models/gemma4-26b
litert-lm import \
  --from-huggingface-repo=litert-community/gemma-4-26B-A4B-it-litert-lm \
  gemma-4-26B-A4B-it-gpu.litertlm \
  gemma4-26b

Note that import points to a specific gemma-4-26B-A4B-it-gpu.litertlm file. Not cosmetic. The README warns that other .litertlm files "may not work well, or not work at all".

Now the minimal agent. It's the same code Google published, no decoration.

# agy_sample.py
import asyncio
import os
from google.antigravity import Agent, LiteRTAgentConfig

MODEL_PATH = os.path.expanduser("~/.litert-lm/models/gemma4-26b/model.litertlm")

async def main():
    print(f"Using local model: {MODEL_PATH}. Inference can take several minutes.")

    config = LiteRTAgentConfig(model_path=MODEL_PATH).lightweight()
    async with Agent(config) as agent:
        response = await agent.chat("What files are in the current directory?")
        async for token in response:
            print(token, end="", flush=True)

if __name__ == "__main__":
    asyncio.run(main())

Three decisions are worth noting. LiteRTAgentConfig(model_path=...) tells the SDK the backend is local, not Gemini. The .lightweight() call keeps the config minimal, which is what you want for a first run on your machine. And async for token in response gives you real streaming: tokens come out as the model generates them, not in one block at the end.

Save the file and run it:

python agy_sample.py

If it's your first run, put the coffee on. Loading 26B parameters into memory takes a while.

The hybrid architect-builder pattern

This is the part that made me stop seeing local models as a toy. Google shipped a demo with an Architect-Builder pattern:

  • A cloud architect (Gemini 3.8 Flash) plans and breaks down the work.
  • A local swarm of Gemma 4 26B does the heavy lifting on the device GPU.

The task: audit and patch three vulnerable modules (auth.py, billing.py, database.py). The numbers from the recorded run:

  • 95 tokens in the cloud. Gemini 3.8 Flash planned from filenames and task descriptions only, without seeing a single line of code.
  • 3,322 tokens (97.2%) executed locally, offline and without calling any API.
  • A green, verified patch, with the proprietary code never leaving the machine.

"97.2% of all tokens (3,322 tokens) run locally and offline without calling a cloud API, delivering fully verified, green patches while keeping proprietary code completely secure on-device." - Source: Google Developers Blog

Think of it as a team. The tech lead (cloud) sets the plan with the bare minimum, and a floor of tireless juniors (local) makes the commits. The lead never saw the repo, only the index. That's why it spends just 95 tokens: it doesn't need more to hand out the work.

The second demo uses the same pattern for a different task. A single prompt, and the local agent writes a Python resource monitor with psutil and rich, generates its requirements.txt, and tests it before handing it over. All on-device.

In that example, the agent doesn't just write code: it writes the dependency file and validates that it runs. That write-verify loop is what separates an assistant from an executor.

Ollama, LM Studio, or vLLM without changing your orchestration

You don't have to stick with Gemma 4 or LiteRT. The SDK accepts any OpenAI-compatible server through LocalOpenAIAgentConfig.

"The Antigravity SDK also offers seamless, plug-and-play support for any OpenAI-compatible server such as Ollama, LM Studio, or vLLM via LocalOpenAIAgentConfig." - Source: Google Developers Blog

What doesn't change matters just as much. Your tools, hooks, and policies stay the same. You swap the backend and that's it. If you already run Ollama at home with a model that fits in 8 GB, you plug it in and your agent, MCP servers and triggers included, keeps working. That decoupling of orchestration from inference is what makes the local ecosystem interchangeable.

What the announcement doesn't tell you

The hardware bar is high. Google recommends at least 24 GB of VRAM or unified memory and a 64K context. If your GPU has 8 GB, forget the 26B and stick with the smaller models in the LiteRT-LM table.

"A4B" isn't size, it's active parameters. Gemma 4 26B A4B is a Mixture of Experts model: 26B total, ~4B active per token. That's why a 26B can be usable on hardware a dense 26B wouldn't survive. I already took that mechanism apart in MoE from the inside: 30B total, 3B active, and it applies here the same way.

Latency on a clean run is published. For Gemma4-E2B on a MacBook Pro M4 Max, LiteRT-LM lists 7,835 tk/s prefill and 160 tk/s decode on GPU, with 0.1 s to first token. On an RTX 4090, 11,234 and 143 tk/s respectively. Mind you: those numbers are for the small models (E2B/E4B), not the 26B.

"A MacBook Pro M4 Max: GPU prefill 7835 tk/s, GPU decode 160 tk/s, Time to First Token 0.1s." - Source: LiteRT-LM Overview

The agent starts in read-only mode. By default it can't write files or run commands. To enable everything (including disk writes) you pass capabilities=CapabilitiesConfig(). That's a sensible safety default: a local agent has access to your filesystem, something a cloud one never has.

Policies are your seatbelt. You can declare which tools the agent may use and which need your approval:

# policies.py
from google.antigravity.hooks.policy import deny, allow, ask_user

policies = [
    deny("*"),                 # Block everything by default
    allow("view_file"),        # Allow reading files
    ask_user("run_command"),   # Ask before executing
]

That leading deny("*") is the pattern I recommend: start closed and open only what you need. It's the equivalent of a firewall for your agent's hands.

Local or cloud? A decision rule

Not everything goes local. My rule, after fighting with both options:

SituationPick
Client code under NDA or complianceLocal
Large, repetitive batch (classify, summarize, extract)Local
No reliable internetLocal
Task that needs the most capable model on the marketCloud
Planning and task decompositionCloud (cheap in tokens)
The best of bothHybrid architect-builder

The trap is wanting 100% on one side. Google's case spends 95 cloud tokens against 3,322 local, and it works because the expensive part is execution, not the plan. Applying that split is exactly what I covered in Model routing and costs: stop sending everything to the expensive model.

And if you're sizing how many concurrent sessions your 24 GB can hold, the calculation that matters is the KV cache, not raw power. I broke that down in 48 KB per token.

Production tips

Five things I'd apply before putting a local agent in a real pipeline:

  1. Sandbox the workspace with workspaces=[WORKING_DIR]. An agent with allow_all() and access to / is a time bomb. Give it one folder and nothing else.
  2. Keep policies closed by default. deny("*") as the base, then open tool by tool as needed.
  3. Pin the context at 64K and don't touch it. The README recommends it. With less, the agent loses the thread on long tasks.
  4. Validate every local patch against regression tests. The gauntlet in the example doesn't approve a patch because it "looks fine", it runs it against a suite. Copy that.
  5. Measure your tokens. If your hybrid spends 60% in the cloud, you got the split wrong. The goal of the pattern is to push execution onto the local GPU.

Wrapping up

What I like most about this announcement is that I can finally choose where each token lives without rewriting my agent. The same Agent, the same tools, the same code. Swap the backend and you're done.

If any of this sounds useful, start small. Install google-antigravity, run the 8-line example, and watch how long Gemma takes to read your directory. Then decide whether the 26B is worth the download.

Original source: Introducing Support for Local AI Models in the Antigravity SDK (Google Developers Blog, September 23, 2026).

Have you run an agent locally, or are you still paying per token in the cloud? Tell me in the comments which model you used and how much VRAM it ate.


> More posts