Skip to content
Comans Services
Menu

Comans Data Foundry · COM-ITOPS-001

Enterprise IT agent training dataset

COM-ITOPS-001 is our first proprietary dataset for enterprise IT agent training and evaluation. It is in development for direct licensing to frontier AI labs and other model developers, with tasks, expert actions and outcomes that can be checked.

A task with a result you can check

The planned dataset combines enterprise IT tasks, records of actions and observations, and outcome checks. Each task brings together the information an agent needs to act and the evidence needed to judge the result.

Objective and starting state

The intended result and the condition of the environment before work begins.

Evidence and constraints

Relevant observations, available tools, permissions and operational limits.

Actions and observations

The steps taken and the information each action produces.

Expected outcome and checks

The intended final state and how success will be assessed.

Provenance and review

How the task was created, reviewed and prepared for its intended use.

Fictional example for explaining the task format

When a successful backup needs a closer look

A backup job reports success, but an attempted restore fails. The task is to investigate the discrepancy and restore the expected test files in an isolated lab.

Evidence

Purpose-created job logs, storage status and a restore procedure.

Constraints

Work within the test environment and preserve the original backup.

What the record captures

The checks selected, what each check reveals and the actions taken as the investigation progresses.

Expected outcome

The intended files are restored to the test location. Their checksums match the reference files, and the original backup remains unchanged.

Follow the evidence and check the result
The fictional task involves investigating the evidence, acting within the test constraints and checking the final state against the expected outcome.
  1. Investigate
  2. Act within the constraints
  3. Verify the final state

A purpose-created example of the proposed task format.

The initial areas we’re exploring

  • Identity and access, including staff changes and permissions.
  • Microsoft 365 and endpoint management.
  • Networks, DNS and service connectivity.
  • Monitoring and incident investigation.
  • Backup, recovery and system dependencies.

These areas guide the first dataset. Each published release will define its actual coverage and limitations.

How quality is being designed in

The planned review process checks whether a task is coherent, its constraints are clear and its expected outcome can be justified. Where practical, a verifier will check system state as well as the agent’s written explanation.

Tasks intended for evaluation need controls over prior exposure. Difficulty labels, task counts and performance results will be published only when there is evidence to support them.

Read our approach to provenance and review

Dataset formats and licensing

Each release will be packaged as a defined dataset with its identifier, version, scope, record types and documentation. Planned formats include structured task records, expert actions and observations, outcome checks and supporting metadata, potentially delivered as JSONL or Parquet with a documented schema.

Buyers will be able to discuss a single dataset licence or regular releases. Record counts, availability, pricing, permitted uses and refresh arrangements will be confirmed for each release.

Discuss your dataset requirements

Tell us about the tasks and record types your team needs, your intended use and whether you want one dataset or regular releases. We’ll discuss fit, development status and future licensing arrangements.

Explore all AI training data work