XOOMAR
Data Refinery

Any public dataset. Cleaned, validated, delivered.

Your analysts lose hours to messy government CSVs: broken encodings, duplicate entities, five date formats in one column, and layouts that drift between releases. We run 31 public datasets around the clock for our own platform. It's the same pipeline can fetch, clean, and ship the public data your team wastes time on. Compliant sources only, no gray-area scraping.

One pipeline, five stages

Source

Gov portals, filings, registries, open APIs, CSV dumps

Polite fetch

Rate-limited, identified, ToS-compliant collection

Clean & normalize

Types, dedupe, encodings, units, entity names

Validate

Schema checks, range rules, row counts vs source

Deliver

One-off or scheduled refresh with monitoring

Raw source

companyratedate
JPMorgan Chase & Co5,25%03/04/26
JP Morgan Chase  5.252026-4-3
jpmorgan-chaseN.A.04/03/2026
Jpmorgan (fullwidth)525bpsApr 3

Delivered

entity_identity_namerate_pctdate
JPMJPMorgan Chase & Co.5.25002026-04-03
4 duplicate rows merged · 1 null documented · source row count reconciled ✓
.csv download.json / JSONLAPI endpointPostgres directcron scheduled refresh

What we take on

Fetch & structure

That agency publishes 400 PDFs a year? We turn them into one table. Parsing, text extraction, format drift handled at the source.

Clean & reconcile

Entity name normalization, dedupe, unit conversion, schema enforcement. Every null documented, never silently dropped.

Keep it fresh

Scheduled refresh with monitoring. You get alerted when the source changes shape, not surprised by silent breakage.

Two ways to engage

One-off delivery

The cleaned dataset delivered once as CSV, JSON, or Excel, with documentation of every transformation and every null. You own the file; the engagement ends when you sign off.

Refresh subscription

We run the pipeline on a schedule, monitor freshness, and deliver via API, Postgres direct, or scheduled file drops. When the source changes shape, we fix it. You never notice.

How we collect

Every fetch runs with a declared user agent and per-source rate limits. We review the source's terms of service before any engagement and respect robots.txt. No personal data, no login-walled content, no gray-area targets. Enterprise buyers shouldn't have to worry about where their data came from.

If a source can't be collected cleanly, we tell you before taking your money.

Proof: we run this pipeline every day

637k+

policy rate rows since 1945

3.8M+

short interest rows

550k+

SEC filings parsed

31

datasets on 24/7 schedules

All public, all live, all collected by this pipeline: browse the free Data Hub →

Describe your dataset

Source, fields, date range, delivery format. A rough sketch is fine. You'll have a fixed quote after we've looked at the source.

Frequently Asked Questions

What sources can you handle?
Any compliant public source: government portals, public registries, regulatory filings, open APIs, and published CSV/Excel dumps. We review the source's terms before quoting. If it can't be collected cleanly, we tell you before taking your money.
One-off or ongoing?
Both. A one-off delivery gets you the cleaned dataset once, as CSV, JSON, or Excel. A refresh subscription means we run the pipeline on a schedule, monitor freshness, and deliver via API, direct to your database, or scheduled file drops.
How is this priced?
Fixed quote after we look at the source, no surprises. A one-off historical extract prices very differently from a monitored refresh feed, so we scope first and quote a number you can plan around.
Can you deliver straight into our database?
Yes: Postgres direct, S3/bucket drops, or an authenticated API endpoint. Scheduled refreshes land in your infrastructure without anyone downloading files by hand.

Data on this page is provided for informational purposes only and is not financial advice. See our editorial policy.