Skip to content

Legacy data format

Prepared train/test datasets used to be flat rows with question, answer, and (for open book QA) context columns. They are now a single messages conversation per example that contains both the question and answer parts, mirroring the chat format the model is trained and served on. This page maps the old format to the new one so you can migrate existing datasets. Datasets in the old format can still be uploaded, but the version is not auto-detected — pass --data-version 0 to distil model upload-data (the data-version query parameter on the API). The default is 1, the messages format described here.

  • One messages array per example. The old question and answer columns are replaced by a conversation: a user turn (the input) and an assistant turn (the expected output).
  • Tool calls use the HuggingFace format. The call lives in the assistant turn’s tool_calls array, and arguments is a real JSON object — not the old stringified answer with a parameters key, and not OpenAI’s stringified arguments.

Question answering, classification, closed book QA

Section titled “Question answering, classification, closed book QA”

Old

{"question": "What is the total amount due?", "answer": "$540"}

New

{"messages": [{"role": "user", "content": "What is the total amount due?"}, {"role": "assistant", "content": "$540"}]}

Old

{"question": "How many students enrolled?", "context": "The university enrolled 5,984 students...", "answer": "5,984"}

New

{"messages": [{"role": "user", "content": "How many students enrolled?"}, {"role": "assistant", "content": "5,984"}], "context": "The university enrolled 5,984 students..."}

The answer string is replaced by an assistant tool_calls array. arguments is a JSON object (HuggingFace format), not a stringified blob, and there is no more parameters key.

Old

{"question": "What's the weather in New York?", "answer": "{\"name\":\"get_weather\",\"parameters\":{\"location\":\"New York, NY\"}}"}

New

{"messages": [{"role": "user", "content": "What's the weather in New York?"}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "get_weather", "arguments": {"location": "New York, NY"}}}]}]}

The conversation history that used to be a stringified JSON array in question, plus the separate answer tool call, are now a single messages array. The target tool call is the final assistant turn.

Old

{"question": "[{\"role\": \"user\", \"content\": \"List files here.\"}, {\"role\": \"assistant\", \"content\": \"\", \"tool_calls\": [{\"type\": \"function\", \"function\": {\"name\": \"ls\", \"arguments\": {}}}]}, {\"role\": \"user\", \"content\": \"Show me config.txt.\"}]", "answer": "{\"name\": \"cat\", \"parameters\": {\"file_name\": \"config.txt\"}}"}

New

{"messages": [{"role": "user", "content": "List files here."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "ls", "arguments": {}}}]}, {"role": "user", "content": "Show me config.txt."}, {"role": "assistant", "tool_calls": [{"type": "function", "function": {"name": "cat", "arguments": {"file_name": "config.txt"}}}]}]}

Note that in the old multi-turn format the assistant turns inside the conversation history already used arguments (a JSON object), while the separate answer used parameters (a stringified object). The new format uses arguments consistently everywhere, in one array, and unifies the last tool call format with the rest of the messages.

For the full per-task specification, see the data preparation overview and the task-specific guides.