Nadim Tuhin
Published on

How I Made Gemini Image Generation Fully Automatic Inside Claude Code

How I Made Gemini Image Generation Fully Automatic Inside Claude Code
Authors
Futuristic AI automation panel

Generated automatically by the system described in this post.


For weeks, generating an image via Gemini inside Claude Code meant a four-step ritual: open Chrome, copy a curl command from DevTools, paste it into the chat, wait for cookies to be parsed. Every two hours, repeat.

This post is about eliminating all of that friction.

The problem: Google's image-gen SIDCC

Gemini's web API gates image generation behind a capability token called SIDCC — a short-lived cookie only issued by Chrome's JavaScript challenge engine during actual image generation. It doesn't appear in SQLite (Chrome stores a weaker version there), it can't be minted by headless browsers, and it can't be transferred between browser profiles.

Every approach I tried confirmed this boundary:

  • Copy Profile 7 cookies → server rejects (device-bound sessions)
  • Playwright with injected cookies → regular SIDCC, blocked
  • CDP / AppleScript → either inaccessible or wrong token type
  • Playwright headless → no SIDCC at all

The image-gen SIDCC is a deliberate security boundary. It cannot simply be bypassed — it has to be worked around architecturally.

The solution: a four-layer fallback pipeline

Rather than fighting the boundary, I built around it with four independent paths:

PriorityLayerMechanismAuth / Prerequisite
1Google Generative Language APIDirect official API callAuto-extracted gcloud keys
2Background Cookie DaemonReal-time Chrome cookie captureSilent daemon sync
3Browser AutomationPlaywright persistent profileOne-time interactive login
4Pollinations.aiPublic generation endpointZero auth, high availability

Each path tries and falls through gracefully. The agent sees a result either way.

Layer 1: Google Generative Language API

The machine already had gcloud credentials configured. Using the API Keys Management API (apikeys.googleapis.com), I automatically retrieved existing keys from the account's gen-lang-client-* AI Studio projects — zero manual key creation needed.

async def _api_generate_image(prompt: str) -> list[str]:
    api_key = os.environ.get("GEMINI_API_KEY", "")
    # Auto-retrieved from local settings if not in env
    model = "gemini-2.5-flash-image"
    url = f"https://generativelanguage.googleapis.com/v1beta/models/{model}:generateContent?key={api_key}"
    # execute request...

When quota allows, this returns a native Gemini image immediately.

A macOS LaunchAgent (com.nadimtuhin.gemini-cookie-daemon) monitors Chrome Profile 7's Cookies SQLite every 300ms. Whenever cookies change, it updates ~/.config/gemini-mcp/latest-cookies.json and settings in real time.

If I generate an image in Chrome normally, the daemon captures the resulting image-gen SIDCC immediately, and the HTTP path starts working automatically for subsequent calls.

Layer 3: Browser automation

A Playwright persistent profile (~/.config/gemini-mcp/browser-profile) provides an isolated browser context with a one-time login. After setup, Chrome handles SIDCC rotation naturally across every generation — no manual cookie management needed.

Layer 4: Pollinations.ai

When local quotas are exhausted and browser profiles are locked, Pollinations.ai provides zero-auth fallback image generation. No API key, no account, no cookies:

async def _pollinations_generate_image(prompt: str) -> list[str]:
    url = f"https://image.pollinations.ai/prompt/{encoded}?nologo=true&model=flux"
    async with cffi.AsyncSession(impersonate="chrome110") as s:
        resp = await s.get(url, timeout=60)
    # save artifact to disk...

What "fully automatic" actually looks like

The pipeline ran for the first time with zero user interaction:

API gen: quota exceeded → falling back
StreamGenerate: blocked (regular SIDCC) → trying Pollinations.ai
[SUCCESS] Saved image to ~/Pictures/gemini/gemini_20260531_051749_0.jpg

Claude Code asked for an image. An image appeared. No curl copying, no DevTools inspection, no manual pasting.

Key takeaways

  1. Google session binding is device-level: Copying cookies between Chrome profiles produces byte-identical values that the server rejects. The device binding is enforced server-side.
  2. The image-gen SIDCC is memory-resident: Chrome stores a standard 74-character SIDCC in SQLite. The image-gen capability token only exists in process memory and only surfaces after an actual image run.
  3. Layered fallbacks beat fragile perfection: The API key has quotas. The daemon depends on Chrome activity. Browser automation needs a profile. Together, they cover 100% of generation requests.
  4. Local credentials are underutilized assets: Existing CLI credentials (gcloud) can negotiate access to internal API keys without prompting the user for secrets.

Architecture files

ComponentResponsibility
scripts/gemini_cookie_daemon.pyPolls Chrome cookies, updates cache
scripts/gemini_browser_imagegen.pyPlaywright UI automation fallback
scripts/gemini_mcp_setup.pyOne-time browser profile initialization
scripts/update_gemini_cookies.pyFallback curl parser to local configuration
com.nadimtuhin.gemini-cookie-daemon.plistAuto-start background daemon

The MCP server chains these in priority order. Claude Code calls gemini_generate_image, receives an artifact path, and continues without friction.