What the importer recognizes
Each line (or array element) is matched against known shapes and normalized to chat format:
Batch output lines (
{"custom_id", "response": {...}}) carry no prompt,
so they are skipped, with the reason counted. Tool calls are preserved
verbatim: agent runs fail at the tool layer (wrong tool, wrong arguments,
ignored result), Omnia’s judges read tool calls, and an import that dropped
them would leave an agent bake-off blind to exactly what distinguishes
quality.
Conversion happens in your browser, before any upload: malformed files
fail fast, and the result shows a per-format breakdown with counted skips
(“974 chat, 3 skipped” reads very differently from “977 prompt/answer”).
Nothing is dropped silently.
Stored answers become the baseline
A trailing assistant message is deliberately kept: it is your incumbent model’s production answer. In an eval against the stored baseline (baseline_model: "__stored__", with sample_filters.dataset_id set to the
imported dataset), the judge grades what your current model actually shipped
against fresh answers from catalog candidates, without Omnia ever calling
your closed model. The incumbent is measured at its production best.
The one-call version, using screening mode:
Grade and calibrate before trusting results
Imported logs carry no grades, so any judge starts uncalibrated. Grade 30 imported exchanges in the review queue and calibrate your judge first; then the bake-off reports corrected rates with intervals instead of raw judge scores. Thirty is the platform floor for calibration to run at all, not a suggestion.Import is a one-shot snapshot, not a pipeline. For continuous capture, pick
one of the integration paths: imported data is enough to
prove whether a switch is worth it, and an integration path is how you keep
measuring after you switch.