Unstructured
Turn PDFs and office documents into clean chunks.
- made by
- Unstructured
Magic Ship is one shop in Vancouver, BC, working remotely with clients worldwide. We are not a partner, reseller, or certified vendor of Unstructured - we just build with this.
What Unstructured is
Unstructured parses documents - PDF, Word, PowerPoint, HTML, email - into typed elements such as titles, narrative text, tables, and list items, with page numbers and coordinates attached. There is an open-source library and a hosted API, including OCR and layout models for scanned pages.
How we use it
Most retrieval failures are ingestion failures, so this is where the early weeks of a knowledge base engagement go. Element types drive chunking: a heading stays attached to the text beneath it, a table is kept whole rather than split mid-row, and page numbers ride along as metadata so an answer can cite a location. Parse quality is checked against a sample of the client's worst documents before any pipeline is built on it.
Where it is the wrong choice
High-fidelity parsing of scanned or heavily formatted PDFs is slow, and on the hosted API it is priced per page, which hurts on a large backfill. Some layouts still need a document-specific parser, and no general tool will save you from writing one.
Service lines it turns up in
Related tools
More in Retrieval and vector
Other tools in the same service lines
Building something on Unstructured?
Send the problem rather than a job spec. You get an answer on scope, on fit, and on whetherUnstructured is even the right call for it.
Start a projectUnstructured and Unstructured are trademarks of their respective owners, used here to say what we work with.