Evaluation and observability

Promptfoo

Config-driven evals and red-teaming for prompts.

made by
Promptfoo

Magic Ship is one shop in Vancouver, BC, working remotely with clients worldwide. We are not a partner, reseller, or certified vendor of Promptfoo - we just build with this.

What Promptfoo is

Promptfoo runs a matrix of prompts, models, and test cases from a config file and scores the outputs with deterministic assertions, similarity checks, or a model acting as judge. Results come back as a side-by-side table in the terminal or a local web view, and it runs in CI. It also has a red-teaming mode for adversarial inputs.

How we use it

This is how a prompt or model change stops being a matter of opinion: the golden question set lives in the repo and a pull request shows what improved and what regressed. Deterministic assertions come first - did it cite a source, did it return valid JSON, did it decline when it should have - and a judge model is used only for what a rule cannot check.

Where it is the wrong choice

Judge-scored evals inherit the judge's biases and cost a model call per case, so a large suite gets slow and expensive and needs sampling. It is also aimed at prompt-level testing; a multi-step agent with tools and state has to be evaluated end to end against a real environment instead.

Building something on Promptfoo?

Send the problem rather than a job spec. You get an answer on scope, on fit, and on whetherPromptfoo is even the right call for it.

Start a project

Promptfoo and Promptfoo are trademarks of their respective owners, used here to say what we work with.