Promptfoo
Config-driven evals and red-teaming for prompts.
- made by
- Promptfoo
- source
- Official Promptfoo site
Magic Ship is one shop in Vancouver, BC, working remotely with clients worldwide. We are not a partner, reseller, or certified vendor of Promptfoo - we just build with this.
What Promptfoo is
Promptfoo runs a matrix of prompts, models, and test cases from a config file and scores the outputs with deterministic assertions, similarity checks, or a model acting as judge. Results come back as a side-by-side table in the terminal or a local web view, and it runs in CI. It also has a red-teaming mode for adversarial inputs.
How we use it
This is how a prompt or model change stops being a matter of opinion: the golden question set lives in the repo and a pull request shows what improved and what regressed. Deterministic assertions come first - did it cite a source, did it return valid JSON, did it decline when it should have - and a judge model is used only for what a rule cannot check.
Where it is the wrong choice
Judge-scored evals inherit the judge's biases and cost a model call per case, so a large suite gets slow and expensive and needs sampling. It is also aimed at prompt-level testing; a multi-step agent with tools and state has to be evaluated end to end against a real environment instead.
Service lines it turns up in
Related tools
More in Evaluation and observability
Other tools in the same service lines
Building something on Promptfoo?
Send the problem rather than a job spec. You get an answer on scope, on fit, and on whetherPromptfoo is even the right call for it.
Start a projectPromptfoo and Promptfoo are trademarks of their respective owners, used here to say what we work with.