Iris vs Langfuse vs Phoenix vs Promptfoo: where each wins, where each loses
Four tools, four different answers to the same question: how do you know what an AI agent did, and whether it was any good? Langfuse is an open-source AI engineering platform, part of ClickHouse since January 2026 ( announcement ). Arize Phoenix is the open-source half of Arize, whose acquisition by Dynatrace was announced in August 2026 ( announcement ). Promptfoo is an open-source CLI for…
Langfuse, Arize Phoenix, Promptfoo, and Iris are four distinct AI evaluation tools, each with unique strengths and weaknesses. Langfuse is an SDK that automatically captures inputs, outputs, timings, and errors of wrapped functions, while Phoenix is an OpenTelemetry platform that instruments applications with auto-instrumentation and exports spans to a collector.
Promptfoo is a CLI for evaluating and red-teaming LLM apps, operating independently of the application. Iris is an MCP server that evaluates agent traces using deterministic rules and publishes precision and recall metrics.
In terms of integration, Langfuse offers numerous SDKs and library integrations, as well as OpenTelemetry compatibility. Phoenix requires OpenTelemetry SDK plus OpenInference auto-instrumentation, with integrations available for popular frameworks. Promptfoo, being a test runner, requires no changes to the app, allowing local or CI-based execution. Iris is integrated into the MCP client configuration, where agents discover it and interact with its tools.
When it comes to evaluation, Langfuse offers LLM-as-a-judge, human annotation, and custom scores via its API and SDK. Phoenix runs evaluators on its server, including LLM-as-a-judge backed by Phoenix-managed prompts and code evaluators with self-hosted backends. Promptfoo's assertion library runs on the agent's machine, analyzing trajectory and tool-call assertions, as well as model-graded rubrics and custom JavaScript or Python.
Iris evaluates agent traces in-process with no model calls and publishes rule precision and recall on a labelled corpus.
Cost-wise, Langfuse offers a free Hobby plan and various pricing tiers for its cloud offering, while self-hosting requires a PostgreSQL, ClickHouse, Redis or Valkey backend, along with S3 or blob storage. Phoenix is free to use locally, with options for Docker, Kubernetes, and PostgreSQL deployments. Promptfoo's Community edition is free forever, with enterprise solutions available upon request. Iris, released under the MIT license, operates as a self-hosted one-process, one-SQLite file deployment.
In summary, Langfuse is an SDK with extensive SDK and library integrations, Phoenix is an OpenTelemetry instrumented platform with server-side evaluation options, Promptfoo is a test runner that directly calls MCP tools, and Iris is an MCP server that evaluates agent traces using deterministic rules. Each tool caters to different aspects of AI agent evaluation, with varying levels of integration, evaluation methods, and costs.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.