LONG-HORIZON TOOL USE

Capture software and technical work for dependable AI.

Capture how engineers investigate systems, choose tools, change artifacts, verify results and recover when an approach fails.

Scope this workflow

From software and technical work activity to a dependable evaluation

A screen recording or final answer does not explain why experienced people make good software and technical work decisions. A useful capture preserves the approved evidence, decision points, exceptions and verified outcome without collecting unrelated activity.

Scope a representative episode

Choose a real software and technical work outcome with a clear start and finish. Define permitted systems, sensitive regions, expected artifacts and stopping conditions before capture so contributors know exactly what is in scope.

Label decisions and exceptions

The valuable evidence is not a list of clicks. Experts explain which facts changed their decision, which policy or standard applied, what uncertainty remained and when a person or a different role needed to take over.

Verify outcomes and reuse failures

Review confirms whether the intended business outcome actually occurred and whether controls were followed. Accepted episodes become demonstrations; corrected or failed cases become regression tests and critical-failure checks for later agent versions.

Calibrate the scoring standard

Reviewers score shared software and technical work cases before production and compare meaningful differences. Calibration clarifies ambiguous instructions, aligns severity thresholds and identifies where the rubric needs another example or a mandatory escalation rule.

Monitor the workflow after release

Evaluation is not finished when an agent passes once. Teams should sample completed work, track failure clusters and add corrected incidents to the regression set. A release gate is valuable only when it remains connected to real outcomes and changing operating conditions.

APPROVED CAPTURE SCOPE

What becomes structured evidence

  • Repository or environment starting state
  • Tool actions, observations and state transitions
  • Verification evidence and recovery paths
  • Final artifacts, checks and deployment boundaries

EXPERT JUDGMENT

Decisions the evaluation must preserve

  • What evidence identifies the real failure?
  • Which change is safe and appropriately scoped?
  • What verification is sufficient before release?

REUSABLE CUSTOMER ASSETS

What your team receives

  • Reproducible technical environments
  • Trajectory and outcome evaluation cases
  • Deterministic checks plus expert rubric
  • Regression comparisons across agent versions

RISK AND REVIEW

What stays explicitly controlled

  • Credential exposure
  • Destructive actions
  • Partial or unverified fixes

Explore other enterprise workflows