Dataset Readiness Review
Inspect authorized local evaluation data for units, missingness, duplicate attempts, task coverage and grouping before calculating engineering comparisons.
5 CANONICAL SKILLS · V0.1.0
Inspect evaluation datasets and analyze complete paired workloads without hiding failures, unknown usage or sampling limits.
Run from a full toolkit checkout. This prints a file manifest, not an installer.
node tools/bundle.mjs resolve engineering-methodsInspect authorized local evaluation data for units, missingness, duplicate attempts, task coverage and grouping before calculating engineering comparisons.
Analyze predeclared engineering comparisons with complete task and attempt records, correct pairing and explicit uncertainty instead of equating fewer tokens with better results.
Evaluate one harness change with a fixed workload, explicit limits, all-attempt accounting, independent integration evidence, and a review decision that preserves unfinished work.
Evaluate whether delegation improves accepted task outcomes using matched inputs, independent checks, repeated sessions, and explicit handoff/integration scoring.
Normalize runtime event captures with explicit identity, usage basis, duplicate handling, missing coverage and independent acceptance boundaries.
Pick the skill needed for the current task. A private adapter supplies project-specific roots, command IDs, accepted contracts, and policy. It should pin and verify the public revision before exposing guidance to an internal agent.
Skill effectiveness remains experimental. Package validation and content hashes do not establish host compatibility, artifact authenticity, or permission to execute.
Read the private-adapter boundary ↗