Developer Tools stdio

Phoenix

Created by Arize-ai

Provides an open-source AI observability platform designed for experimentation, evaluation, and troubleshooting of LLM applications.

How to Connect & Install

Select your client or environment below to copy the verified, ready-to-run installation configuration.

1-line autonomous agent setup instruction. Paste into Claude Code, Codex CLI, or Cursor Agent:
Fetch and configure Model Context Protocol (MCP) server Phoenix from https://www.fastmcp.dev/server/phoenix

Agents automatically discover, install, and execute the server spec without manual token setup.

Connect via direct Stateless MCP v2 / Streamable HTTP endpoint or JSON-RPC:
https://github.com/arize-ai/phoenix

Stateless edge endpoint compliant with MCP 2024-11-05 and v2 draft protocol specs.

Includes server instructions and tool definitions for Claude:
claude mcp add phoenix -- npx -y phoenix

Run in terminal to register server in current project or pass --scope user for global availability.

Configures server instructions and tools for OpenAI Codex & ChatGPT Desktop:
[mcp_servers.phoenix] command = "npx" args = ["-y","phoenix"]

Applies to ChatGPT Desktop app, Codex CLI, and IDE extension across your Codex host.

Add to Cursor project settings (.cursor/mcp.json) or Cursor Settings > Features > MCP:
{ "mcpServers": { "phoenix": { "type": "stdio", "command": "npx", "args": [ "-y", "phoenix" ] } } }

Commit to version control to allow teammates to connect immediately.

Install as an autonomous Agent Skill using the standard skills CLI:
npx skills add https://github.com/arize-ai/phoenix

Standard agent skill installer compatible with Claude Code, Cursor, and agy harnesses.

Server Guidance & Instructions OpenAI & Claude Spec Compliant
Phoenix: Provides an open-source AI observability platform designed for experimentation, evaluation, and troubleshooting of LLM applications. Key capabilities: Leverage LLMs to benchmark application performance using response and retrieval evals.; Track and evaluate changes to prompts, LLMs, and retrieval.

Codex and Claude read this instruction prompt during initialization to orchestrate cross-tool workflows and rate limits.

Top Core Functions
01 Trace LLM application runtime execution
02 Benchmark application performance using evals
03 Manage datasets for fine tuning