AI Chat API

AI APIs for developers

AI Chat API

Build writing, reasoning and coding features with the PoYo AI Chat API. Compare LLM models, conversation protocols, supported inputs, response controls and token pricing.

Matching models

27 models

Kimi K3
Chat

$2.28 1M input

kimi-k3

Moonshot AI's 2.8T-parameter flagship multimodal model with a 1M-token context window, built for long-horizon coding, deep reasoning, knowledge work, and tool-using agents.

Claude Sonnet 5
Chat

$0.85 1M input

claude-sonnet-5

Claude Sonnet 5 API access for agentic coding, long-running agents, browser and computer use, professional workflows, 1M context, and up to 128K output.

DeepSeek V4 Flash
Chat

$0.112 1M input

deepseek-v4-flash

DeepSeek V4 Flash is a fast, economical DeepSeek V4 API model with 1M-context chat, 284B total parameters, and 13B active parameters.

DeepSeek V4 Pro
Chat

$0.348 1M input

deepseek-v4-pro

DeepSeek V4 Pro is a stronger reasoning, coding, and agent workflow model with 1M-context chat, 1.6T total parameters, and 49B active parameters.

GPT-6 Astra
Chat

$8.00 1M input

gpt-6-astra

GPT-6 Astra API for advanced reasoning, coding, and agent workflows. Use Chat Completions or Responses with pay-as-you-go pricing on PoYo.

GPT-6 Sol
Chat

$1.60 1M input

gpt-6-sol

GPT-6 Sol API on PoYo. Use Chat Completions or Responses at $1.6 input and $8 output per million tokens, 20% below official standard rates.

Claude Fable 5.1
Chat

$8.00 1M input

claude-fable-5-1

Anthropic's most capable generally available model for long-horizon coding, knowledge work, research, vision, 1M context, and up to 128K output.

Gemini 3.7 Flash
Chat

$0.06 1M input

gemini-3.7-flash

Google's GA Flash model for agentic coding, multimodal reasoning, web development, long-context knowledge work, and tool-using applications.

Gemini 3.8 Flash
Chat

$0.06 1M input

gemini-3.8-flash

Google's best reasoning and coding Flash model for long-horizon software engineering, autonomous agents, and enterprise workflows.

GPT-5.6
Chat

$0.056 1M input

gpt-5-6-luna

GPT-5.6 API access for Sol, Terra, and Luna through the Responses API, from high-throughput automation to frontier coding and agent workflows.

GPT-6 Luna
Chat

$0.08 1M input

gpt-6-luna

GPT-6 Luna API on PoYo. Use Chat Completions or Responses at $0.08 input and $0.4 output per million tokens, 20% below official standard rates.

Grok 4.6
Chat

$1.60 1M input

grok-4.6

xAI's frontier model for long-running agents, coding, engineering, knowledge work, and ambitious interactive or visual projects.

Grok 4.7
Chat

$1.60 1M input

grok-4.7

Grok 4.7 API on PoYo. Use Chat Completions at $1.6 input and $4.8 output per million tokens, 20% below official standard rates.

Claude Opus 5
Chat

$2.00 1M input

claude-opus-5

Anthropic's flagship agentic model for advanced coding, long-horizon execution, computer use, deep research, and professional knowledge work.

Claude Opus 5.5
Chat

$3.20 1M input

claude-opus-5-5

Claude Opus 5.5 API on PoYo. Use Chat Completions or Claude Messages at $3.2 input and $16 output per million tokens, 20% below official standard rates.

Claude Sonnet 5.5
Chat

$1.60 1M input

claude-sonnet-5-5

Claude Sonnet 5.5 API on PoYo. Use Chat Completions or Claude Messages at $1.6 input and $8 output per million tokens, 20% below official standard rates.

DeepSeek V4.1 Flash
Chat

$0.30 1M input

deepseek-flash

DeepSeek V4.1 Flash is a 552B MoE chat model with native vision, 1M context, 384K max output, and 8B/16B active parameters for cheaper agent workloads.

GPT-6.1 Sol
Chat

$1.60 1M input

gpt-6.1-sol

GPT-6.1 Sol API on PoYo. Use Chat Completions or Responses at $1.6 input and $8 output per million tokens, 20% below official standard rates.

GPT-5.5
Chat

$3.00 1M input

gpt-5.5

GPT-5.5 API access for chat, coding, reasoning, and production agent workflows.

GPT-5.4
Chat

$1.05 1M input

gpt-5.4

GPT-5.4 API access for chat, coding, reasoning, and production agent workflows.

Gemini 3.5 Flash
Chat

$0.90 1M input

gemini-3.5-flash

Gemini 3.5 Flash API access for chat, coding, reasoning, and production agent workflows.

GPT-5.2
Chat

$0.438 1M input

gpt-5.2

GPT-5.2 API access for chat, coding, reasoning, and production agent workflows.

Claude Opus 4.8
Chat

$4.00 1M input

claude-opus-4-8

Claude Opus 4.8 supports long-context chat, agentic coding, professional reasoning, high-output workflows, 1M context, and up to 128K output.

Claude Opus 4.7
Chat

$4.00 1M input

claude-opus-4-7

Anthropic's most capable generally available model for frontier coding, deep reasoning, and high-resolution vision.

Claude 4.6 API
Chat

$1.44 1M input

claude-sonnet-4-6

Claude Sonnet 4.6 and Claude Opus 4.6 API access for coding, agents, computer use, and long-context reasoning.

Claude 4.5 Series
Chat

$0.80 1M input

claude-opus-4-5-20251101

Anthropic's Claude 4.5 series model — Opus, Sonnet & Haiku

Gemini 3 Series
Chat

$0.40 1M input

gemini-3-flash-preview

Google's Gemini 3 series model — Flash Preview & Pro Preview

Chat Model APIs - Pricing and Model Fit

Compare providers, supported inputs, output formats and prices before choosing a model.

Exact task match

Models appear here only when their catalog metadata includes Chat.

Provider comparison

Review model families across providers without leaving the task directory.

API-ready paths

Each model card links to PoYo pricing, examples, playground controls and API details.

ModelProviderTask typesPrice
Kimi K3

kimi-k3

Moonshot AIChat$2.281M input
Claude Sonnet 5

claude-sonnet-5

AnthropicChat$0.851M input
DeepSeek V4 Flash

deepseek-v4-flash

DeepSeekChat$0.1121M input
DeepSeek V4 Pro

deepseek-v4-pro

DeepSeekChat$0.3481M input
GPT-6 Astra

gpt-6-astra

OpenAIChat$8.001M input
GPT-6 Sol

gpt-6-sol

OpenAIChat$1.601M input
Claude Fable 5.1

claude-fable-5-1

AnthropicChat$8.001M input
Gemini 3.7 Flash

gemini-3.7-flash

GoogleChat$0.061M input
Gemini 3.8 Flash

gemini-3.8-flash

GoogleChat$0.061M input
GPT-5.6

gpt-5-6-luna

OpenAIChat$0.0561M input
GPT-6 Luna

gpt-6-luna

OpenAIChat$0.081M input
Grok 4.6

grok-4.6

xAIChat$1.601M input
Grok 4.7

grok-4.7

xAIChat$1.601M input
Claude Opus 5

claude-opus-5

AnthropicChat$2.001M input
Claude Opus 5.5

claude-opus-5-5

AnthropicChat$3.201M input
Claude Sonnet 5.5

claude-sonnet-5-5

AnthropicChat$1.601M input
DeepSeek V4.1 Flash

deepseek-flash

DeepSeekChat$0.301M input
GPT-6.1 Sol

gpt-6.1-sol

OpenAIChat$1.601M input
GPT-5.5

gpt-5.5

OpenAIChat$3.001M input
GPT-5.4

gpt-5.4

OpenAIChat$1.051M input
Gemini 3.5 Flash

gemini-3.5-flash

GoogleChat$0.901M input
GPT-5.2

gpt-5.2

OpenAIChat$0.4381M input
Claude Opus 4.8

claude-opus-4-8

AnthropicChat$4.001M input
Claude Opus 4.7

claude-opus-4-7

AnthropicChat$4.001M input
Claude 4.6 API

claude-sonnet-4-6

AnthropicChat$1.441M input
Claude 4.5 Series

claude-opus-4-5-20251101

AnthropicChat$0.801M input
Gemini 3 Series

gemini-3-flash-preview

GoogleChat$0.401M input

Frequently asked questions

What is an AI chat API?

+

An AI chat API accepts conversation messages and returns a language-model response. It enables writing, reasoning, coding and assistant features in your own application. Available input types and controls depend on the selected model and API protocol.

Which LLM API model should I choose?

+

Evaluate models with the prompts, conversation lengths and outputs your product uses. Compare answer quality, instruction following, speed and token cost. Include difficult and unsuccessful examples in the evaluation rather than selecting from one impressive response.

Does every PoYo chat model use the same request format?

+

No. Use the model documentation to choose OpenAI-compatible Chat Completions, Claude Messages or Gemini native format. Message structures, authentication details and supported controls can differ. Adapt the request to the documented protocol instead of changing only the model name.

How do I preserve multi-turn conversation context?

+

Store the relevant messages in your application and send the history required by the selected protocol. Keep within the model context limit and remove or summarize material that no longer helps the task. A new API request should not be assumed to remember earlier requests automatically.

Does the AI chat API support streaming responses?

+

Check streaming support and the request options for the chosen model and endpoint. If available, handle partial events, completion and errors in your client. Keep a non-streaming path when your workflow needs a complete response before validation or downstream processing.

Can I send images or other media to a chat model?

+

Only use media types and message structures documented for that model. Some LLM endpoints are text-only and others support multimodal inputs. Check size, format and content limitations before implementing uploads, and test how well the model handles your actual media.

Can I request JSON or another structured response?

+

Some models expose structured-output controls; others can only be guided by prompts. Confirm the supported mechanism, supply a clear schema where available and validate the returned data in your application. Do not treat a request for JSON as proof that every response will satisfy your schema.

Can a chat model browse the web or call tools automatically?

+

Only when the endpoint supports the capability and your application provides or enables the required tools. A model can describe a tool call without executing it. Check the documented tool protocol and keep execution, permissions and result handling explicit in your integration.

How can I reduce inaccurate AI chat answers?

+

Provide relevant source material, make the task precise and ask the model to distinguish missing information from known facts. Validate important outputs against trusted data and test recurring failure cases. Model choice and prompting can improve quality, but generated answers still need appropriate review.

How is AI chat API pricing calculated?

+

Chat pricing commonly separates input and output tokens, and some models have additional tier distinctions. Read the selected model pricing and usage response. Include conversation history in your estimate, since repeatedly sending long context can increase the cost of each turn.

What should I do about context errors, rate limits or timeouts?

+

Check the returned error and the selected model limits. Reduce oversized context, control concurrent requests and use bounded retries where appropriate. For a partial streaming response, record what was received and decide how your application should recover instead of silently duplicating a conversation turn.

Where can I find AI chat API integration examples?

+

Open the model page and its API documentation for the correct endpoint, authentication, message body and examples. The model-specific llms.txt resource also identifies the protocol. Keep keys on the server and validate outputs before using them in automated business actions.