wire_api = "responses");而 Codex 内置的 Ollama 供应商仅能访问 localhost,无法连接其他机器上的模型。Tokios 则负责提供 /v1/responses 服务,并将其转换为聊天补全格式供后端调用,如此一来任何运行着 Ollama、llama.cpp、vLLM 或 LM Studio 的远程机器均可作为 Codex 的供应商,且只需通过常规的 https 端点即可实现。
This page assumes you already have Tokios and Codex working together — see Use OpenAI Codex with your own local models for the full config.toml walkthrough. Read on if your model lives on a different machine than the one running Codex, or if you’re troubleshooting a provider that won’t connect.
前置条件
- 在模型所在机器上安装并配置好 连接器——该机器并非运行 Codex 的机器。
- 为该模型设定一个 已注册的部署名称。
- 获取 Tokios 的 API 密钥(
sk-tok-…)。 - Codex CLI installed on the machine where you run
codex
配置 Codex
无论模型位于本地还是网络另一端,供应商配置方式均保持一致;正是 Tokios 确保了地理位置不再成为限制因素。在~/.codex/config.toml 中:
~/.codex/config.toml
gemma-tunnel 替换为您为远程模型所注册的部署名称;该名称并非本地上游模型 ID,也不是 localhost 格式的地址。
为何此方法在本地主机无法奏效
Codex 所使用的自定义提供商必须依赖wire_api = "responses" 才能运作:Codex 会向所有已配置的提供商发送符合 Responses API 规范的请求;而其内置的 Ollama 提供商则被设定为使用固定的 localhost 基础地址,因此无法指向其他机器上的模型。
要让远程模型顺利充当 Codex 的提供商,需同时满足两个条件:该端点必须支持 Responses API,且该端点地址必须是 Codex 在任何运行环境下都能访问的稳定 https 地址。大多数本地推理服务器——如 Ollama、llama.cpp、vLLM 与 LM Studio——仅支持 Chat Completions 协议而非 Responses API。此时 Tokios 网关便位于您的连接器前端,负责接收 Codex 发来的 /v1/responses 请求并将其转换为 Chat Completions 格式后再转发至后端。如此一来,后端完全无需知晓 Responses API 的存在;Codex 也无需配置 localhost 类型的提供商——无论后端位于哪台机器、处于何种网络环境或使用了何种连接器,Codex 始终能通过 https://api.tokios.com/v1 顺利通信。
disable_response_storage = true is required alongside wire_api = "responses". Tokios serves the Responses surface statelessly, so Codex must not rely on server-side response storage between turns.验证功能是否正常
通过该配置执行一条简短的提示指令即可:故障排查
401 —— API 密钥缺失或无效
401 —— API 密钥缺失或无效
Codex isn’t sending a key Tokios recognizes. Confirm
OPENAI_API_KEY is set to your sk-tok-… key in the same shell you launch codex from, and that the key hasn’t been rotated or disabled on the Keys tab.404 — 未找到对应的模型部署
404 — 未找到对应的模型部署
The
model value in [profiles.tokios] doesn’t match a registered deployment. It must be the public name you chose on the Models tab (for example, gemma-tunnel) — not the upstream model id your backend uses locally.503 — 连接器离线或已达负载上限
503 — 连接器离线或已达负载上限
The connector paired with that deployment isn’t currently connected, or it’s at its concurrency cap (
connector_unavailable / connector_busy). Check that the connector process is running on the machine next to the model and shows Online on the Connectors tab.Codex 会转而调用内置的提供商而非 Tokios
Codex 会转而调用内置的提供商而非 Tokios
请确认启动时使用了
--profile tokios 参数。若未指定此参数,Codex 将调用默认提供商,而非 [model_providers.tokios] 中指定的提供商。后续步骤
OpenAI Codex 的基础设置流程
完整的
config.toml 操作指南,包含 OpenAI SDK 的配置方法。快速入门
完成连接器配对并成功注册首个模型。