SkillLift

Type: types/note.md

Evidence basis: source code and attributed README claims at commit 599358b4d4c4a27c0e004df8228ab93026600653, frozen on 2026-09-25. No target rollout or causal experiment was executed. This review covers the flagship portfolio improvement plane and its superproject-owned benchmark adapters; external benchmark, host-agent and provider internals are excluded.

SkillLift improves a task-specific portfolio of skill instructions and auxiliary files through two alternating stages. A planner proposes scoped changes; a bounded editor produces validated patches; Mode A refines branches using a fixed rubric. Benchmark rollouts then evaluate candidates. Mode B revises the rubric to better match those benchmark rankings. The changed portfolio, rubric, hypotheses and scalar outcomes can shape later rounds. This is a builder or improvement plane, with external agents carrying out benchmark tasks. Coordinator — evidenced-by.

The gates have distinct meanings. Portfolio validation checks file scope, hashes, layout and forbidden dependency/environment paths. Verifier scoring asks whether skill text satisfies criteria; malformed output can fall back to a lexical heuristic. Benchmark rewards govern parent promotion, with a higher verifier score allowed to break an equal-reward tie against the parent. None of these checks establishes the truth of an explanatory rationale or success on new tasks. Patch admission, verifier fallback, promotion — evidenced-by.

Three implementation details qualify the README's stronger language. Nondecreasing Mode A score is a surrogate constraint, not a guarantee that task quality cannot fall. Mode B has a finite budget and can retain its best receipt below the requested alignment threshold. The champion's existing reward is reused; it is not rerun for every outer comparison. After adaptation, the coordinator attempts two final trials and reports their maximum valid reward. Alignment selection, final trials — evidenced-by.

Memory is task-local files plus live state. The planner receives current portfolio content, previous plans, direction history and scalar evaluation outcomes. The editor selects files by path and, when needed, by a model's judgment under file/context budgets. Rubric IDs select missing or forbidden behavior feedback for refinement. Controller requests for saved state and cached results are pull; automatic prompt assembly is push. These event-fed plans, patches and rubric revisions establish wired trace learning in the comparison's write-route sense, without establishing successful conjectural learning. Planner context, native editor selection — evidenced-by.

The two benchmark handoffs differ. SkillsBench injects whole skill documents through an owned host patch. WildClaw installs files and tells the task agent to read them. The latter reading is afforded by the handoff, not observed here. Deployment and loaded-skill receipts verify supplied identities and sizes; they do not prove that an agent followed the content. Imported auxiliary assets also prevent a complete representational classification from this boundary. SkillsBench hook, WildClaw loader — evidenced-by.

A conditional recovery mismatch is visible in the code. Mode A can replace a branch's portfolio while keeping its original candidate record. Acceptance stores the candidate ID, and resume rebuilds accepted parents from their saved initial patches. If a content-changing refinement was promoted, replay can reconstruct the earlier draft, fail a later parent-hash check, or conflict with the already frozen refined final tree. Uninterrupted export copies the live refined portfolio. This is a static path finding, not an observed incident. Refinement, replay, final copy — evidenced-by.

The architecture affords content-directed criticism: hypotheses and rubric explanations are explicit, and the revision prompt asks for visible evidence explaining ranking disagreements. Its structural checks require reasons for removed criteria, but cannot validate those reasons. A standing self-improvement pathway is wired at the improvement-plane boundary because benchmark-responsive receipt changes affect its own later scoring and refinement. A narrow reflective path is also wired around its represented assessment criteria and disagreements. Actual valid criticism and an attributable improvement in future capacity remain uninspected. Revision prompt — evidenced-by.

Scope

The repository's reported performance gains remain attributed reports; this analysis did not inspect candidate-linked raw experiments or causal interventions. Remote model weights are not pinned by the inspected client identifiers. External task graders, expected answers, complete host permissions and deployed isolation remain outside the frozen source boundary. Legacy non-portfolio algorithms, tau2 and adoption after export are excluded. Corrected refinement persistence, linked criticism/test traces and suitable interventions would materially change the assessment.

  • Exact analysis — see-also: canonical routes, source excerpts, both lenses and the fourteen-axis memory profile
  • Conjectural learning — defined-in: distinguishes criticism-linked capacity improvement from a trace-fed write route
  • Reflective system — defined-in: the narrow causally connected self-representation claim
  • Self-improving system — defined-in: the dispositional improvement pathway and its boundary