Inside the
data refinery.
Datology is the platform for data curation. With Datology, your team runs a frontier-level data process to clean, curate, create synthetic data, and compose the data for training runs, all in your environment.
Clean. Curate. Create. Compose.
Four research-driven stages turn your public, proprietary, and licensed data into multi-stage training datasets. All running in your environment, operated by your team.
Clean
Remove malformed, empty, or short documents that don’t provide any signal for model training. Then ensure you decontaminated against evaluation benchmarks. All running at petabyte scale. What's left is data your model can learn from.
Connect customer sources; unify schema, repair encoding
Remove degenerate samples based on length, symbol-to-word ratio, etc.
N-gram leakage check across all training sources
Curate
Tailor the data by quality, task, and taxonomy to your use case. The pipeline understands what's in the corpus, keeps what your model needs, and shapes the training distribution around the specific job your model has to do.
Heuristic and learned quality and taxonomy classifiers
N-gram and embedding-based sample similarity rejection
Adjust the sampling distribution to emphasize high-quality data
Reshape the corpus to your use-cases, using examples you provide
Create
Generate synthetic data to maximize task-relevance and diversity. Real data has coverage gaps the world doesn't fill on its own. The synthetic data fills them without the failure modes that break naive synthetic approaches. Proven at trillion-token scale in production.
Multiple signals, including quality and task-relevance, are used to select documents for rephrasing
Rephrasing approaches avoid the failure modes of de novo generation
Carefully-tuned rephrasing methodology maximizes diversity and the mileage of your data
Compose
Mix data sources optimally to produce staged training datasets. Different phases of training need different data: pre-training wants broad coverage, mid-training wants domain focus, and annealing wants quality signal. The pipeline decides the composition, so each stage gets what it needs.
Principled, quantitative mixture design
Stage-aligned datasets for mid-training and annealing
Runs in your environment.
Datology deploys in your cloud. Your data is curated in place and never leaves your environment.
Deploys in your VPC
Datology runs in your cloud account, so you have complete control.
Your data never leaves
Curation happens in place. Nothing is copied out of your environment.
A model you own
The result is a model that you own, not rent

Your data
YOUR STORAGE
Data curation pipeline
RUNS IN YOUR VPC

Staged training datasets
ready to train
Learn more
Real results from teams building their own models.
Legal domain adaptation.
Mid-training against a proprietary legal corpus broke through the post-training ceiling on both public and proprietary evals, on a budget under 1% of base pre-training (100B mid-training vs. 15T pre-training).
Frontier open-weights model.
A frontier-class open weights MoE built without a frontier budget or research team. ~10 people, 20T tokens curated by Datology. Competitive with models trained by 170-person teams that raised $2B.

Build your next model on better data.
Meet with a data curation expert to see how data quality can be a game changer for your business.
