Training examples
Questions, answers, and worked examples for a defined task. Each example includes the context a model needs to answer correctly.
The outputTask, context, and reference answer
Data preparation & evaluation
Datarmon is a boutique data company. We're developing task-specific datasets to help AI teams train models and check their answers.
What we're buildingSource
“Disconnect power before removing the cover. Wait at least 60 seconds before servicing.”
Sample equipment instruction
Question
Can I remove the cover while the power is still connected?
A specific task, not a general topic
Expected answer
No. Disconnect the power before removing the cover.
An answer supported by the source
The answer must say to disconnect power before removing the cover. It must not invent an exception or treat the 60-second wait as permission to leave power connected. Keeping the source alongside the answer makes that check possible.
Our work
Our private beta focuses on three types of data for teams adapting models to specialized tasks.
Questions, answers, and worked examples for a defined task. Each example includes the context a model needs to answer correctly.
The outputTask, context, and reference answer
Two responses to the same question, a judgment about which is better, and a written explanation of that judgment.
The outputPaired responses and review notes
Cases kept separate from training data, with expected answers and clear scoring rules. Used to check where a model still gets things wrong.
The outputTest cases and scoring criteria
How we approach it
Set the scope, source material, and rules for judging an answer. Identify what the model should do when information is missing.
Check examples against the agreed rules. Resolve ambiguous cases before producing more of the same data.
Keep the examples, source references, review notes, and known limitations together so the dataset can be inspected and reused.
About Datarmon
We're building Datarmon for AI teams that need data for a particular kind of work, rather than a general-purpose dataset. Our focus is on preparing examples and tests that can be checked against clear requirements.
The work described here is the focus of our private beta. There is no public dataset catalog or self-serve product at this stage.