Evaluation and observability

Playwright

Browser automation for tests and for real verification.

made by
Microsoft

Magic Ship is one shop in Vancouver, BC, working remotely with clients worldwide. We are not a partner, reseller, or certified vendor of Microsoft - we just build with this.

What Playwright is

Playwright drives Chromium, Firefox, and WebKit from one API with automatic waiting, network interception, tracing, and parallel execution. It runs headless in CI, records a trace and video of failures, and can generate a test by recording a flow.

How we use it

We use it to prove a change works in the running product rather than only in a unit test, which is the bar we hold ourselves to before calling anything done. The trace and screenshots from a failed run get attached to the ticket, so a reviewer sees the failure instead of reading a description of it. It is also how an agent verifies its own frontend change rather than asserting success.

Where it is the wrong choice

Browser tests are the slowest and flakiest layer in any suite, and a team that pushes assertions down here instead of into unit and API tests ends up with a suite nobody trusts. It also cannot tell you whether a model's answer was any good, only that the page rendered.

Building something on Playwright?

Send the problem rather than a job spec. You get an answer on scope, on fit, and on whetherPlaywright is even the right call for it.

Start a project

Playwright and Microsoft are trademarks of their respective owners, used here to say what we work with.