Model icon
gemini-3.7-flash
Chat
Model:
Gemini 3.7 Flash is Google's GA multimodal workhorse for agentic coding, web development, and 1M-context knowledge workflows.
Chat

This Playground is for demo purposes. Data is only valid in the current window and will be cleared on refresh.

Type a message...

Configuration
1.0
1.0
Pricing details

Transparent pricing with no hidden fees. Pay as you go.

Googlegemini-3.7-flash
Input
PoYo price
$0.600/1M tokens
120 credits
Official price
$0.750/1M tokens
Official
You save
20%
Googlegemini-3.7-flash
Cached input
PoYo price
$0.060/1M tokens
12 credits
Official price
$0.075/1M tokens
Official
You save
20%
Googlegemini-3.7-flash
Output (including thinking tokens)
PoYo price
$3.00/1M tokens
600 credits
Official price
$3.75/1M tokens
Official
You save
20%

* Actual fees are based on the final output.

Available on PoYo

Complete guide to using Gemini 3.7 Flash API for Fast, Capable Agents

Gemini 3.7 Flash API for Fast, Capable Agents

Gemini 3.7 Flash is Google's generally available Flash workhorse for agentic coding, web development, multimodal reasoning, and complex knowledge work. It supports a 1,048,576-token input context and up to 65,536 output tokens.

Use gemini-3.7-flash through /v1/chat/completions or the Gemini-native /v1beta/models path. PoYo charges 120 input, 12 cached-input, and 600 output credits per million tokens—20% below Google's introductory Standard rates.

Key Features of Gemini 3.7 Flash

Official Google demonstrations show the model moving from planning to finished, interactive results across coding, visual, physical, and document tasks.

01

Build Interactive Games from an Idea

Gemini 3.7 Flash can plan and implement interactive experiences while keeping the visual result and code behavior aligned.

  • Multi-step coding and debugging
  • Visual feedback during iteration
  • Fast Flash latency for repeated tool loops
Official Gemini 3.7 Flash game-building demonstration

02

Create Higher-Fidelity Web Interfaces

Stronger design adherence helps turn visual references into polished landing pages and audit existing interfaces for implementation gaps.

  • Design-to-code workflows
  • Responsive frontend generation
  • Visual auditing and refinement
Official Gemini 3.7 Flash interactive landing-page demonstration

03

Coordinate Tools for Physical Tasks

Agentic planning and multimodal understanding let the model interpret a scene, choose actions, and coordinate tools over several steps.

  • Image and video understanding
  • Disciplined multi-step planning
  • Function calling and tool orchestration
Official Gemini 3.7 Flash robotics demonstration

04

Turn Long Documents into Useful Outputs

A million-token context window supports extended document analysis, evidence synthesis, calculations, and report-oriented workflows.

  • PDF and long-document input
  • Structured extraction and analysis
  • Detailed text output up to 64K tokens
Official Gemini 3.7 Flash annual-report analysis demonstration

What Can You Build with Gemini 3.7 Flash?

01

Agentic Software Engineering

Plan and execute multi-step coding work, navigate repositories, call tools, debug failures, and verify changes with fewer retries.

02

Web and UI Development

Turn design references into functional interfaces, audit existing frontends, and iterate with visual feedback.

03

Document and Knowledge Work

Reason across long reports, PDFs, research material, and enterprise documents within a 1M-token context.

04

Multimodal Analysis

Combine text, images, video, audio, and PDFs for review, extraction, understanding, and structured answers.

05

Tool-Using Agents

Build agents with function calling, code execution, search grounding, structured output, and URL context.

06

Adjustable Reasoning

Choose low, medium, or high thinking levels to balance response speed, cost, and task depth.

Gemini 3.7 Flash Benchmark Comparison

Selected August 2026 benchmarks compare Gemini 3.7 Flash with Gemini 3.6 Flash, Claude Sonnet 5, and GPT-5.6 Terra across coding, agentic execution, web development, document work, and long-context reasoning.

BenchmarkGemini 3.7 FlashGemini 3.6 FlashClaude Sonnet 5GPT-5.6 TerraNotes
Artificial Analysis Intelligence Index56525557Composite intelligence
FrontierCode 1.1 Main43.6%34.4%42.7%41.3%Production code quality
DeepSWE v1.165.3%48.6–49.0%53.8%69.6%Long-horizon software engineering
WebDev Arena (Elo)1588153815411523Web development
Terminal-Bench 2.185.8%78.0%80.4%87.4%Agentic terminal coding
Terminal-Bench 3.014.9%5.4%14.6%20.8%General agent capabilities
AutomationBench30.4%17.0%10.7%23.6%Enterprise workflow automation
GDP.pdf34.0%22.0%28.0%24.7%Complex document processing
Harvey LAB-AA (Legal)90.7%85.1%90.1%85.2%Complex legal workflows
OSWorld-2.047.9%33.8%—50.2%Agentic computer use
GDM-MRCR v2 (128k)97.0%91.8%81.5%93.5%Long-context performance
HLE-Verified53.6%51.2%31.0%51.1%Multidisciplinary expert reasoning

Selected benchmark results from the supplied Gemini 3.7 Flash knowledge base, current as of August 2026. Scores use each benchmark's published evaluation setup.

How to Use Gemini 3.7 Flash on PoYo

1. Create an API key. Sign in and create a PoYo key.

2. Choose a format. Use POST /v1/chat/completions for OpenAI-compatible clients, or POST /v1beta/models/gemini-3.7-flash:generateContent for Gemini-native payloads. Streaming is available through :streamGenerateContent.

3. Set the model. Send gemini-3.7-flash as the model ID. Read the Gemini API documentation.

Gemini 3.7 Flash API Frequently Asked Questions

Is Gemini 3.7 Flash available on PoYo now?

Yes. Use the gemini-3.7-flash model ID with either supported endpoint format.

Which endpoints are supported?

Use /v1/chat/completions for OpenAI-compatible requests or /v1beta/models/gemini-3.7-flash:generateContent for Gemini-native requests. Both non-streaming and streaming flows are supported.

How much does Gemini 3.7 Flash cost?

PoYo charges 120 credits per million input tokens, 12 credits per million cached-input tokens, and 600 credits per million output tokens, including thinking tokens.

What media can the model understand?

Google lists text, image, video, audio, and PDF input with text output. Native image generation, audio generation, and Live API output are not supported by this model.

How large is the context window?

The model supports up to 1,048,576 input tokens and 65,536 output tokens.

Which thinking levels are available?

Gemini 3.7 Flash supports low, medium, and high thinking levels. The minimal level is not supported.

Does cached input include storage charges?

No. The 12-credit rate applies to cached input tokens read by a request. Any separate context-cache storage duration charge is outside that rate.

How is Gemini 3.7 Flash different from 3.6 Flash?

Gemini 3.7 Flash is an iterative reasoning upgrade with substantial gains in production coding, long-horizon software engineering, web development, agent reliability, and document understanding while retaining the same 1M context and multimodal inputs.

What is the knowledge cutoff?

Google lists March 2026 for most domains, with some areas potentially limited to January 2025. Use Search Grounding or retrieval for current information.

Does Gemini 3.7 Flash support Computer Use or fine-tuning?

Computer Use is available as a Preview feature. Fine-tuning is not supported.

How does it compare with Claude Sonnet 5 and GPT-5.6 Terra?

The selected benchmarks show Gemini 3.7 Flash leading several web, automation, document, legal, long-context, and expert-reasoning evaluations while remaining competitive on coding and agent tasks. Results vary by benchmark, so use the comparison table rather than one aggregate claim.

When does Google's introductory pricing end?

Google's introductory Standard rates end December 31, 2026. Google lists higher Standard rates beginning January 1, 2027; PoYo pricing should be reviewed before that date.

Why Use Gemini 3.7 Flash on PoYo

01

Two API Formats

Use OpenAI-compatible Chat Completions or Gemini-native GenerateContent with one model ID.

02

One PoYo Key

Evaluate and ship Gemini alongside other models without managing another integration credential.

03

Transparent 20% Savings

See input, cached-input, and output rates in both credits and US-dollar equivalents.

04

Playground and Monitoring

Test both request formats, review service status, and track usage before moving to production.