AI 聊天 API

面向开发者的 AI API

AI 聊天 API

通过 PoYo AI 聊天 API 接入写作、推理和代码能力。对比 LLM 模型的调用协议、输入支持、响应控制和 token 价格。

匹配模型

27 个模型

Kimi K3
聊天

$2.28 100 万输入

kimi-k3

Moonshot AI's 2.8T-parameter flagship multimodal model with a 1M-token context window, built for long-horizon coding, deep reasoning, knowledge work, and tool-using agents.

Claude Sonnet 5
聊天

$0.85 100 万输入

claude-sonnet-5

Claude Sonnet 5 API access for agentic coding, long-running agents, browser and computer use, professional workflows, 1M context, and up to 128K output.

DeepSeek V4 Flash
聊天

$0.112 100 万输入

deepseek-v4-flash

DeepSeek V4 Flash is a fast, economical DeepSeek V4 API model with 1M-context chat, 284B total parameters, and 13B active parameters.

DeepSeek V4 Pro
聊天

$0.348 100 万输入

deepseek-v4-pro

DeepSeek V4 Pro is a stronger reasoning, coding, and agent workflow model with 1M-context chat, 1.6T total parameters, and 49B active parameters.

GPT-6 Astra
聊天

$8.00 100 万输入

gpt-6-astra

GPT-6 Astra API for advanced reasoning, coding, and agent workflows. Use Chat Completions or Responses with pay-as-you-go pricing on PoYo.

GPT-6 Sol
聊天

$1.60 100 万输入

gpt-6-sol

GPT-6 Sol API on PoYo. Use Chat Completions or Responses at $1.6 input and $8 output per million tokens, 20% below official standard rates.

Claude Fable 5.1
聊天

$8.00 100 万输入

claude-fable-5-1

Anthropic's most capable generally available model for long-horizon coding, knowledge work, research, vision, 1M context, and up to 128K output.

Gemini 3.7 Flash
聊天

$0.06 100 万输入

gemini-3.7-flash

Google's GA Flash model for agentic coding, multimodal reasoning, web development, long-context knowledge work, and tool-using applications.

Gemini 3.8 Flash
聊天

$0.06 100 万输入

gemini-3.8-flash

Google's best reasoning and coding Flash model for long-horizon software engineering, autonomous agents, and enterprise workflows.

GPT-5.6
聊天

$0.056 100 万输入

gpt-5-6-luna

GPT-5.6 API access for Sol, Terra, and Luna through the Responses API, from high-throughput automation to frontier coding and agent workflows.

GPT-6 Luna
聊天

$0.08 100 万输入

gpt-6-luna

GPT-6 Luna API on PoYo. Use Chat Completions or Responses at $0.08 input and $0.4 output per million tokens, 20% below official standard rates.

Grok 4.6
聊天

$1.60 100 万输入

grok-4.6

xAI's frontier model for long-running agents, coding, engineering, knowledge work, and ambitious interactive or visual projects.

Grok 4.7
聊天

$1.60 100 万输入

grok-4.7

Grok 4.7 API on PoYo. Use Chat Completions at $1.6 input and $4.8 output per million tokens, 20% below official standard rates.

Claude Opus 5
聊天

$2.00 100 万输入

claude-opus-5

Anthropic's flagship agentic model for advanced coding, long-horizon execution, computer use, deep research, and professional knowledge work.

Claude Opus 5.5
聊天

$3.20 100 万输入

claude-opus-5-5

Claude Opus 5.5 API on PoYo. Use Chat Completions or Claude Messages at $3.2 input and $16 output per million tokens, 20% below official standard rates.

Claude Sonnet 5.5
聊天

$1.60 100 万输入

claude-sonnet-5-5

Claude Sonnet 5.5 API on PoYo. Use Chat Completions or Claude Messages at $1.6 input and $8 output per million tokens, 20% below official standard rates.

DeepSeek V4.1 Flash
聊天

$0.30 100 万输入

deepseek-flash

DeepSeek V4.1 Flash is a 552B MoE chat model with native vision, 1M context, 384K max output, and 8B/16B active parameters for cheaper agent workloads.

GPT-6.1 Sol
聊天

$1.60 100 万输入

gpt-6.1-sol

GPT-6.1 Sol API on PoYo. Use Chat Completions or Responses at $1.6 input and $8 output per million tokens, 20% below official standard rates.

GPT-5.5
聊天

$3.00 100 万输入

gpt-5.5

GPT-5.5 API access for chat, coding, reasoning, and production agent workflows.

GPT-5.4
聊天

$1.05 100 万输入

gpt-5.4

GPT-5.4 API access for chat, coding, reasoning, and production agent workflows.

Gemini 3.5 Flash
聊天

$0.90 100 万输入

gemini-3.5-flash

Gemini 3.5 Flash API access for chat, coding, reasoning, and production agent workflows.

GPT-5.2
聊天

$0.438 100 万输入

gpt-5.2

GPT-5.2 API access for chat, coding, reasoning, and production agent workflows.

Claude Opus 4.8
聊天

$4.00 100 万输入

claude-opus-4-8

Claude Opus 4.8 supports long-context chat, agentic coding, professional reasoning, high-output workflows, 1M context, and up to 128K output.

Claude Opus 4.7
聊天

$4.00 100 万输入

claude-opus-4-7

Anthropic's most capable generally available model for frontier coding, deep reasoning, and high-resolution vision.

Claude 4.6 API
聊天

$1.44 100 万输入

claude-sonnet-4-6

Claude Sonnet 4.6 and Claude Opus 4.6 API access for coding, agents, computer use, and long-context reasoning.

Claude 4.5 Series
聊天

$0.80 100 万输入

claude-opus-4-5-20251101

Anthropic's Claude 4.5 series model — Opus, Sonnet & Haiku

Gemini 3 Series
聊天

$0.40 100 万输入

gemini-3-flash-preview

Google's Gemini 3 series model — Flash Preview & Pro Preview

聊天模型 API - 价格和模型适配

选择模型前,对比供应商、支持的输入、输出格式和价格。

精确任务匹配

只有模型目录元数据包含 Chat 的模型才会出现在这里。

供应商对比

无需离开任务目录,即可查看不同供应商的模型系列。

API 就绪路径

每张模型卡都会跳转到 PoYo 的价格、示例、Playground 控件和 API 详情。

模型供应商任务类型价格
Kimi K3

kimi-k3

Moonshot AIChat$2.28100 万输入
Claude Sonnet 5

claude-sonnet-5

AnthropicChat$0.85100 万输入
DeepSeek V4 Flash

deepseek-v4-flash

DeepSeekChat$0.112100 万输入
DeepSeek V4 Pro

deepseek-v4-pro

DeepSeekChat$0.348100 万输入
GPT-6 Astra

gpt-6-astra

OpenAIChat$8.00100 万输入
GPT-6 Sol

gpt-6-sol

OpenAIChat$1.60100 万输入
Claude Fable 5.1

claude-fable-5-1

AnthropicChat$8.00100 万输入
Gemini 3.7 Flash

gemini-3.7-flash

GoogleChat$0.06100 万输入
Gemini 3.8 Flash

gemini-3.8-flash

GoogleChat$0.06100 万输入
GPT-5.6

gpt-5-6-luna

OpenAIChat$0.056100 万输入
GPT-6 Luna

gpt-6-luna

OpenAIChat$0.08100 万输入
Grok 4.6

grok-4.6

xAIChat$1.60100 万输入
Grok 4.7

grok-4.7

xAIChat$1.60100 万输入
Claude Opus 5

claude-opus-5

AnthropicChat$2.00100 万输入
Claude Opus 5.5

claude-opus-5-5

AnthropicChat$3.20100 万输入
Claude Sonnet 5.5

claude-sonnet-5-5

AnthropicChat$1.60100 万输入
DeepSeek V4.1 Flash

deepseek-flash

DeepSeekChat$0.30100 万输入
GPT-6.1 Sol

gpt-6.1-sol

OpenAIChat$1.60100 万输入
GPT-5.5

gpt-5.5

OpenAIChat$3.00100 万输入
GPT-5.4

gpt-5.4

OpenAIChat$1.05100 万输入
Gemini 3.5 Flash

gemini-3.5-flash

GoogleChat$0.90100 万输入
GPT-5.2

gpt-5.2

OpenAIChat$0.438100 万输入
Claude Opus 4.8

claude-opus-4-8

AnthropicChat$4.00100 万输入
Claude Opus 4.7

claude-opus-4-7

AnthropicChat$4.00100 万输入
Claude 4.6 API

claude-sonnet-4-6

AnthropicChat$1.44100 万输入
Claude 4.5 Series

claude-opus-4-5-20251101

AnthropicChat$0.80100 万输入
Gemini 3 Series

gemini-3-flash-preview

GoogleChat$0.40100 万输入

常见问题

什么是 AI 聊天 API?

+

AI 聊天 API 接收对话消息并返回语言模型响应,可用于应用中的写作、推理、编程和助手功能。支持的输入类型与控制取决于模型和调用协议。

应该选择哪个 LLM API 模型?

+

用产品实际提示词、对话长度和输出要求测试,比较质量、指令遵循、速度与 token 成本。评估中应包含困难和失败样例,不要只根据一次出色回答选型。

所有 PoYo 聊天模型都使用相同请求格式吗?

+

不是。按模型文档选择 OpenAI 兼容 Chat Completions、Claude Messages 或 Gemini 原生格式。消息结构、鉴权和支持的参数可能不同,切换时不应只改模型名称。

如何保留多轮对话上下文?

+

在应用中保存相关消息,并按协议发送所需历史。控制上下文长度,对无关材料进行删除或总结;不能假定新请求会自动记住之前的调用。

AI 聊天 API 支持流式输出吗?

+

核对所选模型与接口的流式选项。支持时应处理增量事件、完成和错误;需要先验证完整结果再进入后续流程的场景,可保留非流式调用方式。

可以向聊天模型发送图片或其他媒体吗?

+

只使用模型文档支持的媒体类型和消息结构。部分 LLM 为纯文本,其他模型支持多模态;接入上传前检查格式、大小和内容限制,并测试真实素材。

可以要求返回 JSON 等结构化内容吗?

+

部分模型提供结构化输出控制,其他模型主要依赖提示词引导。确认机制并在应用端验证数据;要求返回 JSON,并不代表每次响应都一定满足你的 schema。

聊天模型会自动联网或执行工具吗?

+

只有接口支持且应用提供或启用相应工具时才能实现。模型描述工具调用不代表已经执行,接入时应明确工具协议、执行权限与结果处理。

如何减少错误回答?

+

提供相关来源、明确任务,并要求区分缺失信息与已知事实。对重要结果使用可信数据验证,持续测试重复失败的样例;模型和提示词优化不能代替必要审核。

AI 聊天 API 如何计算价格?

+

通常分别计算输入与输出 token,部分模型还有其他档位。结合当前价格和使用量响应估算,并计入反复发送的对话历史,长上下文会增加每轮成本。

上下文超限、限流或超时怎么办?

+

检查错误与模型限制,减少过长上下文、控制并发并按情况设置有界重试。流式响应只收到部分内容时,应记录状态并设计恢复方式,避免静默重复一个对话轮次。

在哪里查看 AI 聊天 API 接入示例?

+

模型页、API 文档和模型专属 llms.txt 提供对应端点、鉴权、消息结构与示例。将 Key 保存在服务端,并在结果进入自动化业务操作前进行验证。