Evaluation and observability

Langfuse

Tracing and eval storage for LLM applications.

made by
Langfuse

Magic Ship is one shop in Vancouver, BC, working remotely with clients worldwide. We are not a partner, reseller, or certified vendor of Langfuse - we just build with this.

What Langfuse is

Langfuse records traces of LLM calls - prompts, responses, tool calls, latency, token counts, cost - and groups them into sessions and users. Datasets and scores live next to the traces, so an eval run and the production traffic it is meant to model sit in one place. It is open source and can be self-hosted.

How we use it

We instrument an agent before we tune it, because without traces a prompt change is guesswork with a confident tone. Failing production traces get promoted into a dataset, and that dataset becomes the regression suite the next change has to pass. Self-hosting is the usual choice when prompts or retrieved passages are confidential.

Where it is the wrong choice

It only sees what you send it, so an application already instrumented with OpenTelemetry ends up with two partial views of the same request. In a shop with a mature observability stack, exporting spans to the existing backend is the better fit.

Building something on Langfuse?

Send the problem rather than a job spec. You get an answer on scope, on fit, and on whetherLangfuse is even the right call for it.

Start a project

Langfuse and Langfuse are trademarks of their respective owners, used here to say what we work with.