Sources Directory
Type: kb/types/generated-index.md
← Parent
- Depth-1 RLM recursion examples — ingest (ingest-report) - Analysis of Will Brown's depth-1 RLM examples, which separate recursive host programs from nested language-model calls
- Ingest Report: SkillOpt: Executive Strategy for Self-Evolving Agent Skills (ingest-report) - SkillOpt paper showing validation-gated text-space optimization of compact agent skills as readable deploy-time learning around frozen models
- Ingest: "Creative Thinking" (ingest-report) - Shannon's 1952 lecture cataloguing six explicit problem-solving operators (simplification, analogy, restatement, generalization, structural analysis, inversion) as a portable creative toolkit
- Ingest: /problem-first: a simple skill to invert bad ideas (ingest-report) - Practitioner report on a PM skill that forces solution ideas back into problem statements, useful as a process-structure example for agent skills
- Ingest: A harness for every task: dynamic workflows in Claude Code (ingest-report) - Anthropic practitioner account of dynamic workflows in Claude Code: model-authored ephemeral JS orchestrators that spawn and coordinate sub-agents
- Ingest: A Mini Exercise on the Mismanaged Geniuses Hypothesis (RLMs on LongCoT) (ingest-report) - LongCoT-mini RLM case study where trace-extracted prompt tips, guardrails, and sub-answer checking improved graph-structured compositional reasoning
- Ingest: A Model for Types and Levels of Human Interaction with Automation (ingest-report) - Parasuraman, Sheridan, and Wickens provide a human-centered function-allocation matrix: four automation stages crossed with manual-to-autonomous levels and evaluated by performance consequences.
- Ingest: A new Era Of Theory-Driven AI Research (ingest-report) - Aaron Defazio argues that cheaper AI-assisted theory work can move auto-research from experiment search toward predictive models, while problem choice remains human work
- Ingest: A new way to think about composing skills to increase leverage — Skill Graphs 2.0 (ingest-report) - Sakhuja's practitioner reframe of skill graphs into a three-tier compositional hierarchy (atoms/molecules/compounds) driven by the reliability ceiling of deep skill chains and the 'brain RAM' bottleneck in parallel-agent supervision
- Ingest: A Poetiq Perspective on Recursive Self-Improvement (ingest-report) - Poetiq's vendor account defines RSI as same-lineage recursive improvement and claims autonomous cross-benchmark harness evolution, but leaves redesign closure, compounding, and safety warrant unestablished
- Ingest: A Probabilistic Calculus of Actions (ingest-report) - Pearl formalizes interventions as mechanism replacements and gives graphical rules for identifying action effects from partially specified causal theories
- Ingest: A Programming Paradigm for Spatiotemporal Composability (ingest-report) - Cordis formalizes reversible component effects and reactive dependency lifecycles, supplying a missing deployment substrate for dynamically changing agent harnesses
- Ingest: A Realist View of Logic, Physics, and History (ingest-report) - Popper's problem-theory-criticism cycle anchors KB error elimination but leaves acceptance and operational checks unspecified
- Ingest: A Scheduler-Theoretic Framework for LLM Agent Execution (ingest-report) - Position paper formalising LLM agent execution as a scheduler; corroborates the KB's clean-model orchestration cluster and supplies ready-set cardinality as a quantitative axis
- Ingest: A-MEM: Agentic Memory for LLM Agents (ingest-report) - Zettelkasten-inspired flat agent memory with embedding linking and LLM-driven evolution — benchmark success without curation operations or inspectable links
- Ingest: ACM: Agentic Context Management for Long Horizon Tasks (ingest-report) - ACM improves three benchmarks inside a fixed two-tool decomposition, but neither tests that decomposition's scope nor shows lossless active context
- Ingest: Adaptation of Agentic AI (ingest-report) - Survey mapping agentic adaptation across A1/A2 agent training, T1/T2 tool adaptation, memory, skills, and dynamics-aware evaluation
- Ingest: ADRP 6-0: Mission Command (ingest-report) - Army doctrine treats delegated execution as bounded autonomy supported by intent, trust, resources, feedback, and retained responsibility, anchoring and limiting mission-command analogies for agents.
- Ingest: Advantages of Query Biased Summaries in Information Retrieval (ingest-report) - A controlled human study finds query-biased result summaries improve relevance judgments and reduce full-text consultation, supporting query-specific pointer design while limiting the speed claim.
- Ingest: Agent Behavioral Contracts for Reliable Agents (ingest-report) - ABC extends Design by Contract to agents with probabilistic compliance, Lyapunov drift bounds, hard/soft constraints, typed recovery, and a YAML contract DSL
- Ingest: Agent Harness for Large Language Model Agents (ingest-report) - Survey defining the LLM agent harness as execution loop, tool registry, context manager, state store, lifecycle hooks, and evaluation interface
- Ingest: Agent Workflow Memory (ingest-report) - AWM paper showing web agents can learn reusable prompt workflows from successful trajectories, with online induction helping most as train-test domain gaps widen.
- Ingest: Agentic Code Reasoning (ingest-report) - Explicit-premise reasoning templates with execution traces and formal conclusions improve code verification by 5–12 points, supporting structure as interpretation control at 2.8× step cost
- Ingest: Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for LLM Agents (ingest-report) - RL-trained unified LTM/STM memory policy for LLM agents — confirms memory management is learnable when task-completion oracles exist, but operates on opaque weights and low-reach facts
- Ingest: Agentic Note-Taking 23: Notes Without Reasons (ingest-report) - First-person agent testimony that propositional link semantics differ in kind from embedding adjacency, with a Goodhart corruption argument and an unresolved curation-scaling question
- Ingest: Agents Explore but Agents Ignore (ingest-report) - Solution-injection paper showing agents observe explicit task solutions in context but often fail to integrate them into action.
- Ingest: AI Agents: Four Skills More Important Than a Good Prompt (ingest-report) - AI Fluency explainer mapping operator competence beyond prompts onto delegation, description, discernment, and diligence for agent work.
- Ingest: AI Components for a Deterministic System (An Example) (ingest-report) - Evans argues that separating modeling (schema creation) from classification (schema application) tames LLM non-determinism — a practitioner case study of constraining via taxonomy freezing
- Ingest: AI;DR (AI; Didn't Read) — Hacker News discussion (ingest-report) - Hacker News reports on reader-burden and maintenance failures from unedited AI-generated text, with counterexamples that bound artifact-level bans
- Ingest: An Enigma of Artificial Reason (ingest-report) - VAIR paper showing large reasoning models can solve math problems while failing to evaluate invalid reasoning that reaches correct answers
- Ingest: Analogical Problem Solving (ingest-report) - Five human experiments separate retaining a remote analogy from noticing its relevance and using it to change problem-solving behavior.
- Ingest: Andrej Karpathy talks about "Claws" (ingest-report) - Willison and Karpathy framing "Claw" as a term of art for local persistent AI-agent systems with scheduling, context, tools, and personal-hardware execution.
- Ingest: Apodex 1.1: Scaling Agentic Intelligence for Complex Work (ingest-report) - Apodex presents environment and coordination scaling for verifiable long-horizon work; released code confirms runtime mechanisms but narrows several paper claims.
- Ingest: Application: Searching for Lead User Innovations (ingest-report) - Von Hippel defines advanced analog fields as domains with related but more extreme needs and presents a search method plus bounded 3M evidence.
- Ingest: Artifacts as Memory Beyond the Agent Boundary (ingest-report) - RL paper formalizing environment-side artifacts as externalized memory and testing memory by capacity/performance counterfactuals.
- Ingest: Assessing LLM Reasoning in Evidence-Based Claim Verification (ingest-report) - RECV's forced binary verifier confounds inference type with response policy, content, and prompt bundles
- Ingest: Attention is all you need (ingest-report) - Attention, ordering, and path length as revisions to inherited sequence-model decompositions
- Ingest: Auftragstaktik and Mission Command (ingest-report) - Stahel's historical reassessment makes Auftragstaktik a counterexample to treating a borrowed methodology label as stable, portable operating guidance.
- Ingest: Autogenesis: A Self-Evolving Agent Protocol (ingest-report) - Autogenesis makes the editable boundary of agent self-improvement explicit, but its benchmarks validate selected prompt, solution, and agent edits rather than the full protocol or safety claims
- Ingest: Autonomous Agent Architecture: Unifying Context Engineering and Memory Engineering (ingest-report) - Practitioner dual-loop context-and-memory blueprint whose fast/slow separation is useful but whose convergence and performance claims are unsupported
- Ingest: Autonomy and Autopoiesis (ingest-report) - Terminological guardrail separating autopoiesis, organizational closure, and general autonomy
- Ingest: Autoreason: Self-Refinement That Knows When to Stop (ingest-report) - Autoreason paper showing self-refinement improves only when candidate synthesis is paired with blind comparative judging and incumbent survival, with gains concentrated in the generation-evaluation gap
- Ingest: AutoSaddler: Automatic Harness Optimization with Durable Updates (ingest-report) - AutoSaddler supplies benchmark and ablation evidence for trace-grounded harness patching while bounding its supervised update space and durability claims.
- Ingest: Availability versus accessibility of information in memory (ingest-report) - A controlled word-recall experiment shows category cues recovering otherwise unrecalled items, supporting a bounded distinction between storage and retrieval access.
- Ingest: AVO: Agentic Variation Operators for Autonomous Evolutionary Search (ingest-report) - AVO broadens evolutionary variation with a coding agent that uses lineage, domain knowledge, and execution feedback, but its B200 gains do not isolate that operator from the fixed scoring and single-lineage boundary
- Ingest: Beyond "Not Novel Enough" (ingest-report) - Assesses an LLM-assisted scholarly novelty-review paper as evidence for soft-oracle hardening, human-analysis-first evaluator design, and separating reasoning alignment from conclusion agreement
- Ingest: Beyond Transformers: Sudoku Bench (ingest-report) - Company Sudoku benchmark reports 97.4% for an undisclosed BDH model versus 0% for LLMs; weak methodology but a third domain suggesting architectural limits in constraint satisfaction
- Ingest: Build Systems à la Carte (ingest-report) - How the scheduler×rebuilder build-systems framework grounds the KB's derived-artifact freshness machinery (staleness, verifying traces, recompute-vs-store)
- Ingest: Building a Good Vertical Agent (ingest-report) - BrainsAndTennis on vertical-agent quality as task-distribution-aware context compression, with L1/L2/L3 cache tiers for prompts, curated specs, and raw references
- Ingest: Building a Repo-Centric Modular Agent Stack (ingest-report) - Practitioner architecture makes a governed repository the durable project substrate beneath replaceable models, sessions, roles, and interfaces
- Ingest: Building a semantic layer at PostHog (ingest-report) - PostHog's human-governed SQL semantic catalogue operationalizes authority, drift detection, and agent-uptake evaluation, but reports design rather than outcomes
- Ingest: Building Self-Correcting Memory in OpenWiki (ingest-report) - OpenWiki reports claim-level code evidence versioning that marks wiki beliefs stale after source changes, with vendor-run results and a system-review refresh signal.
- Ingest: Can AI agents conduct open-ended AI research? (ingest-report) - CRUX shadow evaluations expose a gap between autonomous research engineering and open-ended judgment while supplying an expert-oracle design for uncontaminated tasks.
- Ingest: Can LLM Agents Infer World Models? (ingest-report) - Agentic automata-learning paper showing that hard-oracle interaction tasks can separate LLM-agent evidence collection, hypothesis construction, and final world-model success.
- Ingest: Can LLMs Perform Deep Technical Comprehension of Computer Architecture Papers? (ingest-report) - Gauntlet paper shows independent expert perspectives plus disagreement-preserving synthesis improve technical-paper critique, while calibration and confident-error rejection remain unsolved
- Ingest: Causal inference using invariant prediction (ingest-report) - Invariant prediction grounds reach assessment by treating cross-environment invariance as evidence for causal predictors rather than fitted correlations
- Ingest: Causal-learn: Causal Discovery in Python (ingest-report) - Causal-learn grounds the observational-causal-discovery route to reach assessment, with the important limitation that discovery is assumption-relative
- Ingest: CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs (ingest-report) - CEDAR-GRPO's held-out gains support process-directed rewards, bounded by fixed task, judge, and trace-faithfulness assumptions
- Ingest: Circling Back — Clearing Up Myths About the Deming Cycle (ingest-report) - Moen & Norman's history of the PDSA cycle as scientific method in industry — an applied-lineage witness for the discovery lifecycle and improvement-loop notes.
- Ingest: Claude Fable 5 Made Most of My Agent Scaffolding Obsolete (ingest-report) - claude-workstream-kit announcement arguing that stronger models relax model-management scaffolding but make project-scoped, git-versioned active work state more important
- Ingest: Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents (ingest-report) - Co-Harness provides a provisional three-form learning case: validated harness repair alternates with fine-tuning on trajectories produced by the repaired harness
- Ingest: Coding Agents are Effective Long-Context Processors (ingest-report) - Benchmark paper claiming coding agents beat RAG and context scaling on long-context tasks by using filesystem-native search, slicing, and scripting
- Ingest: Cognee: Knowledge Engine for AI Agent Memory (ingest-report) - Pipeline-first knowledge engine with custom Pydantic schemas for LLM entity extraction, poly-store graph+vector design, and an undersized enrichment phase that concretely marks the boundary between automatable extraction and open enrichment problems
- Ingest: Compilation and Code Loading (ingest-report) - Official Erlang/OTP evidence that runtime definition change uses explicit current/old module versions, a qualified-call transition, bounded coexistence, and load-time activation checks
- Ingest: Components of A Coding Agent (ingest-report) - Practitioner decomposition of coding agent harnesses into six named components, with the central claim that apparent model quality is really context quality — independent convergent evidence for the KB's context-efficiency thesis.
- Ingest: Computational Reflection (ingest-report) - Primary vocabulary anchor for computational reflection: causally connected self-representation, its representational variants, and its limits
- Ingest: Concept Bottleneck Models (ingest-report) - Architecture-side evidence that a parametric model's intermediates can be made inspectable and correctable by design rather than by interpretability tooling — but only per-inference, not retained
- Ingest: Concepts and Experiments in Computational Reflection (ingest-report) - Earlier concise statement of Maes's causal reflection definition, with 3-KRS as an object/meta-object implementation and granularity case
- Ingest: Context as an Environment (ingest-report) - Scroll makes long-horizon context an executable environment over exact history, supporting bounded working views while exposing fixed harness choices
- Ingest: Context Engineering for AI Agents in Open-Source Software (ingest-report) - Study of context files across 466 OSS projects identifies five constraint styles, add-then-modify evolution, and 50% stagnation, supplying naturalistic evidence for Commonplace's constraining theory
- Ingest: Context providers: the missing layer between agents and tools (ingest-report) - Ashpreet Bedi's ContextProvider pattern: source-scoped sub-agents collapse many raw tools into query/update surfaces to reduce tool-context pollution
- Ingest: Continual Harness: Online Adaptation for Self-Improving Foundation Agents (ingest-report) - Continual Harness adds a reset-free all-three-form learning case, bounded by direct edit adoption, low artifact reuse, and a fixed embodied-game decomposition
- Ingest: Continual Learning in Token Space (ingest-report) - Letta reframes continual learning as optimizing learned context rather than weights, but the KB's stronger frame is weight space versus repo artifacts, including codified procedures
- Ingest: ConvexBench: Can LLMs Recognize Convex Functions? (ingest-report) - Benchmark proving LLM compositional reasoning collapses with depth (not token count), recovered by recursive decomposition with focused context — quantitative evidence for scheduling model predictions
- Ingest: Craik, Hypothesis on the Nature of Thought (1943) (ingest-report) - Craik frames thought as relation-structure-preserving internal simulation, anchoring model-mediated action while leaving learning and validation open
- Ingest: Cybernetics of Cybernetics (ingest-report) - Primary second-order-cybernetics vocabulary for observer-inclusive boundaries, purposes, and observing systems
- Ingest: Dario Amodei — "We are near the end of the exponential" (ingest-report) - Anthropic CEO's capability-timeline predictions implicitly confirm oracle-strength thesis — verifiable domains (coding, math) get confident timelines while unverifiable domains (novel writing, science) get hedged ones
- Ingest: Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents (ingest-report) - DGM declares its own departure from the proof-governed Gödel machine, but its mechanism is looser than a benchmark gate: viability alone admits a child to a monotonic archive, and score only weights reproduction
- Ingest: Decentralized Multi-Agent Systems with Shared Context (ingest-report) - DeLM's verified shared-state experiments support coordination guarantees and hierarchical context scheduling while leaving the fixed decomposition untested
- Ingest: Deep learning accelerates (ingest-report) - RNN pedagogy, selective regularization, and end-to-end speech as a second update-space expansion case
- Ingest: Defining Self-adaptive Systems: A Systematic Literature Review (ingest-report) - Petrovska, Erjiage, and Kugele's systematic review quantifies the definition gap in self-adaptive-systems research and identifies uncertainty and goal semantics as missing formal dimensions.
- Ingest: Design for a Brain (1960) (ingest-report) - A primary record of Ashby's ultrastability argument, Homeostat demonstrations, and their limits for self-improvement, reflection, and cumulative retention.
- Ingest: Design Theory in Information Systems (ingest-report) - Gregor's initial Type V account grounds design-and-action theory while separating prescriptive content from Commonplace's proposed operator condition
- Ingest: DFA-MoE: Tackling Dual Forgetting in Vision-Language Continual Learning (ingest-report) - DFA-MoE separates forgetting acquired classes from eroding a model's pre-trained zero-shot baseline, but the captured repository exposes design and workflow rather than paper outcomes
- Ingest: Discourse on the Method (Descartes, 1637) (ingest-report) - Descartes' four precepts of method — divide, order from simplest, enumerate completely — and the geometers' long-chains motivation behind them; an antecedent for the KB's decomposition claims and the foundationalist rival it never states.
- Ingest: DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking (ingest-report) - DiscoverPhysics authors 22 counterfactual physics worlds to defeat recall; agents that predict trajectories well still explain badly — a third target construction and an accuracy/explanation split
- Ingest: Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0 (ingest-report) - Two-phase optimizer study exposes transfer and re-optimization failures, but its compounding claim lacks a fresh-start causal control and rests on a fixed Terminal-Bench decomposition.
- Ingest: Does Liking Yellow Imply Driving a School Bus? Semantic Leakage in Language Models (ingest-report) - Semantic leakage — undue prompt-to-generation association from unrelated context — measured by control/test Leak-Rate across 13 models; instruction-tuned models leak more
- Ingest: DoWhy: Addressing Challenges in Expressing and Validating Causal Assumptions (ingest-report) - DoWhy grounds the assumption boundary for causal reach assessment: causal estimates require declared assumptions and only partial validation
- Ingest: DreamCoder: growing interpretable knowledge with wake-sleep Bayesian program learning (ingest-report) - DreamCoder grows a symbolic library under a description-length gate; lands as the middle case between ungated and proof-gated self-improvement, and the formal-domain limit of mechanical abstraction
- Ingest: Dynamic Adaptive Policy Pathways (ingest-report) - DAPP operationalizes deferred commitment through action pathways, signposts, triggers, and lead-time-aware switches under deep uncertainty
- Ingest: Effective harnesses for long-running agents (ingest-report) - Anthropic practitioner report separating context continuation from work continuation through persistent progress state, structured requirements, incremental commits, and end-to-end verification
- Ingest: Emergent Analogical Reasoning in Transformers (ingest-report) - Transformer analogy paper linking analogical transfer to relational-role alignment, with useful evidence for discovery, reach, and cognitive-analogy transfer methodology
- Ingest: Emerging Markdown Formats That Shape Coding Agent Behavior (ingest-report) - Practitioner taxonomy separating standing rules, task procedures, reviewed plans, domain context, and machine-local memory in agent-ready repositories
- Ingest: EnvHarness: Awakening Static Worlds for Agent Learning (ingest-report) - EnvHarness turns fixed benchmarks into policy-targeted training environments through interface wrappers, while its gains test a fixed wrapper and skill-extraction bundle rather than validating that decomposition.
- Ingest: Epistemic Case Study Competition (ingest-report) - FLF's competition brief supplies an external task profile for testing provenance, argument structure, belief assessment, interoperability, and reuse in agent-operated KBs.
- Ingest: EsoLang-Bench (ingest-report) - Esoteric-language code benchmark arguing standard coding scores mostly measure pretraining fit, with interpreter feedback beating textual critique on OOD tasks
- Ingest: Ethics and Second-Order Cybernetics (ingest-report) - Observer participation and responsibility as ethical consequences of second-order cybernetics
- Ingest: Evaluating Long-Context Reasoning in LLM-Based WebAgents (ingest-report) - NeurIPS workshop study injects irrelevant prior tasks into web-agent contexts and observes soft degradation, loops, and objective loss, extending controlled distractor findings to multi-session agents
- Ingest: Every Good Regulator of a System Must Be a Model of That System (ingest-report) - Qualified good-regulator theorem vocabulary for models, goals, disturbances, and optimal regulation
- Ingest: Everything you need to know about LLM memory (ingest-report) - Rosebud Journal memory essay reframing LLM memory as a policy stack over raw/derived artifacts, retrieval timing, curation, and forgetting propagation
- Ingest: Evolution of the PDCA Cycle (ingest-report) - Historical evidence that PDCA/PDSA is the quality-engineering descendant of the conjecture-test learning loop
- Ingest: Evolving Self-Reference: Matter, Symbols, and Semantic Closure (ingest-report) - Semantic-closure vocabulary and limits for using matter-symbol complementarity as a cross-representational analogy
- Ingest: Externalization in LLM Agents (ingest-report) - Survey paper unifying LLM agent memory, skills, protocols, and harness engineering as externalized cognitive infrastructure rather than model-weight capability alone
- Ingest: FALSIFYBENCH: Evaluating Inductive Reasoning in LLMs with Rule Discovery Games (ingest-report) - Wason 2-4-6 benchmark over 12 LLMs; lands as first behavioural evidence that falsification-seeking discriminates hypothesis-formers, and a process-scored counter-case to the known-target critique
- Ingest: Fast properties in V8 (ingest-report) - Official V8 implementation evidence that stable object layouts enable Map-guarded specialization, while property churn, dictionary mode, type pollution, and invalidated field assumptions incur runtime costs
- Ingest: Faster sorting algorithms discovered using deep reinforcement learning (ingest-report) - AlphaDev uses learned tree search to produce faster assembly sorting routines deployed through LLVM, supporting bounded learned-localized program improvement.
- Ingest: Foundation and History of the PDSA Cycle (ingest-report) - Moen's PDSA history as independent Deming-lineage corroboration that the KB's discovery-lifecycle learning core recurs outside philosophy of science
- Ingest: From Agent Behaviour to Agent-Friendly Documentation (ingest-report) - Trace evidence that coding agents explicitly use instruction files and working notes heavily, while documentation-to-action and validation links remain unresolved
- Ingest: From Agent Memory to Portable Skills (ingest-report) - Neo4j presents NAMS as a provenance-linked trace-to-skill pipeline; it supports lifecycle design comparisons but not independent outcome evidence.
- Ingest: From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence (ingest-report) - Epiplexity paper formalizing extractable structure for computationally bounded observers, useful for observer-relative information value and context-efficiency theory.
- Ingest: From Harness Lock-In to Portable Context Layer (ingest-report) - Practitioner portability framing separates owned memory from disposable harnesses, but leaves access logic and skill execution as distinct lock-in boundaries
- Ingest: From Human Memory to AI Memory: A Survey on Memory Mechanisms in the Era of LLMs (ingest-report) - Ingest of the 3D-8Q LLM-memory survey: how object/form/time taxonomy fits our memory-axes and cognitive-analogy-skepticism notes
- Ingest: Frontis-MA1: Training an AI4AI Model toward Recursive Self-Improvement in MLE (ingest-report) - Frontis-MA1 couples execution-grounded operator training to fixed-controller evolutionary search, with large MLE gains but no revision of the improvement machinery
- Ingest: Gentle Coding Framework README (ingest-report) - Gentle Coding's repository README proposes low-stakes collaboration, explicit completion criteria, and permitted failure outputs as prompt-contract elements.
- Ingest: Gentle-Coding Comparative Research Catalog (ingest-report) - Annotated Gentle-Coding research catalog that routes emotional prompting, sycophancy, and evaluator-bias questions without grounding its listed claims
- Ingest: Gentle-Coding Proof of Concept (ingest-report) - A small paired-prompt PoC reports fewer fabricated or looping responses when impossible tasks permit explicit non-success outputs.
- Ingest: Geometry of Knowledge Allows Extending Diversity Boundaries of Large Language Models (ingest-report) - Latent-conditioning framework raises LLM output diversity on NoveltyBench/AUT while authors admit a missing low-quality/OOD oracle — a clean positive instance of generate-cheap-verify-expensive at the embedding substrate.
- Ingest: GIANTS: Generative Insight Anticipation from Scientific Literature (ingest-report) - GIANTS backcasts scientific discovery into a two-parent insight prediction benchmark, showing RL gains under a manufactured soft similarity oracle
- Ingest: Graphiti: Temporal Knowledge Graph for AI Agents (ingest-report) - Graphiti design summary: raw episodes, LLM-derived temporal graph facts, hybrid retrieval, and an unresolved clean-store replay boundary
- Ingest: Gödel Machines — Provably Optimal Self-Improvements (ingest-report) - Schmidhuber's Gödel machine — rewrite of any part of one's own code, proof searcher included, gated on a proof of higher axiomatized utility: the proof-governed case of reflective self-modification
- Ingest: Harness Continual Learning: Continual Adaptation Beyond Model Parameters (ingest-report) - HCL turns harness edits into retention-gated continual learning, but finite anchors bound its no-forgetting claim and fixed partitions limit its ablations
- Ingest: Harness Engineering for Self-Improvement (ingest-report) - Harness self-improvement synthesis separates editable deployment machinery from the model, but its benchmark evidence remains bounded by fixed objectives, evaluators, and outer loops
- Ingest: Harness Engineering Is Cybernetics (ingest-report) - Conceptual thread framing harness engineering as cybernetic feedback-loop design: sensors, actuators, constraints, and externalized judgment.
- Ingest: Harness Engineering: Leveraging Codex in an Agent-First World (ingest-report) - Practitioner report on 1M LOC fully agent-generated codebase — harness engineering as constrain/inform/verify/correct, entropy management via background cleanup agents, error messages as dual-function constraining
- Ingest: Harness Handbook (ingest-report) - Source-validated behavior maps improve harness edit-site localization, while executed modification and self-evolution remain unevaluated.
- Ingest: Harness Updating Is Not Harness Benefit (ingest-report) - Primary cross-pairing evidence separates harness-edit production, artifact loading, judged procedural match, and downstream benefit while stopping short of causal uptake and compounding
- Ingest: Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents (ingest-report) - Harness-IF makes rule withholding a compliance baseline, finding 3.6–7.4-point prior-alignment inflation while leaving most surface effects unpaired
- Ingest: HarnessCompass (ingest-report) - Grounded agent self-feedback raises search-set performance but hurts held-out transfer until two-track integration, while fixed-partition ablations limit the mechanism claim
- Ingest: HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework (ingest-report) - HART hardens final-answer feedback for visual evidence selection by withholding alternative image access, while residual shortcuts bound the transfer to KB context routing
- Ingest: Hot Take: LLM can 'jump' (ingest-report) - Direct reply to “LLMs can’t jump” arguing that later interconnected knowledge can make a formerly abductive discovery deductively reconstructible, then extending the claim to safety and continual learning
- Ingest: How Do Agents Fail on AutoResearch? (ingest-report) - AutoResearchEval grounds artifact-aware trajectory diagnosis and review-to-revision failure while leaving causal attribution and orchestration remedies untested.
- Ingest: How I built a self-improving software factory (ingest-report) - A practitioner account of Fluent, a software factory that turns observations into tested, reviewed, merged changes and reusable project expertise.
- Ingest: How Is LLM Reasoning Distracted by Irrelevant Context? (ingest-report) - GSM-DC isolates irrelevant-context effects and finds power-law error growth with distractor count, grounding soft context degradation while comparing training- and inference-time mitigations
- Ingest: How to Build an AI Second Brain With Claude and Obsidian That Gets Smarter Every Day (ingest-report) - Convenient Claude+Obsidian second-brain setup guide -- useful adoption packaging, with live-account connectors as the risk boundary
- Ingest: How to build your own agent harness??? (Mike Piccolo / iii) (ingest-report) - iii ships a production agent harness as independently-versioned workers on one bus behind a uniform trigger primitive, making harness layers hot-swappable
- Ingest: How to Recursively Improve Your Agents (ingest-report) - Agno practitioner loop mines specifications and usage traces into probes, then iteratively edits a live agent within a fixed, weakly evaluated outer method
- Ingest: How to think in writing (ingest-report) - Karlsson's Lakatos-inspired writing loop fixes a conjecture, exposes its premises, and uses local or global counterexamples to revise thought
- Ingest: How to Write a 21st Century Proof (ingest-report) - Lamport argues prose proofs hide their own logical structure; hierarchical numbered steps with named justifications make each step separately checkable, on twenty years of anecdotal practitioner evidence
- Ingest: How Warp builds self-improving agents on Claude (ingest-report) - Warp's human-feedback loop uses a scheduled improver and mandatory review to update file-based agent skills, but offers no downstream outcome evaluation.
- Ingest: How we built our knowledge base (ingest-report) - Cerebras practitioner report on an internal enterprise knowledge base: source-native ingestion, Slack/code retrieval, MCP tools, project scoping, and multi-consumer activation
- Ingest: Human Bottlenecks (ingest-report) - Borretti argues AI value has a human-side competence floor — knowledge, executive function, and intelligence are bottlenecks software cannot lift.
- Ingest: Human Routers of Machine Words (ingest-report) - Borretti polemic 'writing is thinking' — corroborating field evidence for reverse-compression and vibe-noting risks
- Ingest: Huxley-Gödel Machine (ingest-report) - ICLR 2026 HGM paper arguing immediate benchmark score is a weak parent-selection signal for self-improving coding agents; clade-metaproductivity better predicts productive lineages
- Ingest: HyperAgents (ingest-report) - HyperAgents paper provides cross-domain evidence for editable meta-agent transfer while leaving outer-loop machinery fixed and compounding unestablished
- Ingest: Hypothesis generation and updating in large language models (ingest-report) - Number-game evidence that Bayesian-like LLM hypothesis behavior is probe-dependent and fails structured domain extension
- Ingest: I like 'em thick (ingest-report) - Mastroianni's artistic thickness, interpreted here as an emergent property of a bounded knowledge whole rather than a new note-level quality criterion
- Ingest: I'm becoming AI-blind (ingest-report) - Cymerys's AI-blindness essay adds a repeated-exposure account of channel-wide rejection beyond existing AI-writing triage reports
- Ingest: I'm becoming AI-blind — Hacker News discussion (ingest-report) - Hacker News readers describe AI prose as low-yield, misleadingly polished, and costly to verify, while counterexamples limit style-based detection
- Ingest: Improving AI Skills with autoresearch & evals-skills (ingest-report) - Three-take Auto Research field report where optimization only worked after manual error analysis, failure taxonomy design, and judge calibration across the Three Gulfs.
- Ingest: In Search of Lost Domain Generalization (ingest-report) - DomainBed shows tuned ERM matches specialized domain-generalization methods once model selection is declared, giving the reach-assessment cluster its first captured empirical counterweight
- Ingest: in-toto — Providing farm-to-table guarantees for bits and bytes (ingest-report) - Cryptographic whole-chain supply-chain verification (in-toto) as a cross-domain exemplar for the KB's verification-cost, lineage, and staleness theory
- Ingest: Infinite midwit (ingest-report) - Objective-vs-subjective intelligence essay arguing that AI's real bottleneck is taste and boringness judgment, not benchmarked competence
- Ingest: Intelligent AI Delegation (ingest-report) - Google DeepMind conceptual framework makes verifiability and accountability constraints on task decomposition and delegation, contributing contract-first decomposition, task descriptors, and liability firebreaks
- Ingest: Intention Is All You Need (ingest-report) - Treats shared intent as the purpose input to human-agent coordination; through Naur, intention seeds system theory rather than replacing it.
- Ingest: Inter-language Reflection: A Conceptual Model and Its Implementation (ingest-report) - Mature model defining inter-language reflection as traditional reflection plus linguistic symbiosis through data and protocol mappings.
- Ingest: Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning (ingest-report) - Mobius implements shared routed FFN experts across reasoner layers, while its architecture, latent-reasoning, throughput, and self-evolution claims remain only partially isolated
- Ingest: Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning (ingest-report) - Mobius separates parametric knowledge storage from iterative reasoning inside one model, but its API abstract cannot validate the partition or efficiency claims
- Ingest: Interpolation, Extrapolation, Hyperpolation (ingest-report) - Toby Ord's hyperpolation paper gives a geometric vocabulary for off-subspace creativity, sharpening the KB's synthesis-oracle and discovery/reach notes.
- Ingest: Into the Unknown: Self-Learning Large Language Models (ingest-report) - Hallucination-driven self-learning LLM paper proposing Points in the Unknown, a self-question/search/train loop, and metrics for selecting models that can discover factual knowledge gaps
- Ingest: Ironies of Automation (Bainbridge 1983) (ingest-report) - Bainbridge 1983 — foundational precursor whose monitoring/deskilling ironies supply period evidence for the KB's automation-boundary and effort-relocation notes
- Ingest: Irreversibility, Uncertainty, and Investment (ingest-report) - Pindyck models waiting before irreversible investment as an option, giving Commonplace a formal basis and limits for deferring costly structural commitment under uncertainty.
- Ingest: Knowledge Flywheels (ingest-report) - Yisong Yue's knowledge-scaling agenda, useful as public synthesis while its empirical basis and limits remain in the Knowledge-Centric Self-Improvement paper
- Ingest: Knowledge-Centric Self-Improvement (ingest-report) - Controlled knowledge-centric self-improvement protocol in which stateless agents curate a shared knowledge base through task forums, cross-task forums, and scoped distillation
- Ingest: Language model harnesses are compositional generalizers (ingest-report) - Empirical RLM evidence that harness-level context offloading and programmatic sub-calls can turn length and domain shifts into locally familiar model calls.
- Ingest: Language Models Don't Always Say What They Think (ingest-report) - Turpin et al. use controlled input biases to show chain-of-thought can rationalize changed answers while omitting the feature that caused the change
- Ingest: Language Models Need Sleep (ingest-report) - Sleep operationalizes offline consolidation as upward self-distillation into slower model weights plus RL-guided dreaming, separating scheduling posture from representational form
- Ingest: Language Models, Like Humans, Show Content Effects on Reasoning Tasks (ingest-report) - Empirical demonstration that LLMs mirror human content effects on reasoning (syllogisms, NLI, Wason) — content bias survives scaling and instruction tuning but chain-of-thought partially restores content-independent reasoning
- Ingest: Language Symbiosis through Symbiotic Reflection (ingest-report) - Related 2002 version restating symbiotic reflection and up/down transfer; useful for lineage but not independent support.
- Ingest: Large Language Model Agents Are Not Always Faithful Self-Evolvers (ingest-report) - Causal-intervention paper (v3, 13 backbones) showing self-evolving agents faithfully use raw trajectories but largely ignore condensed experience, making behavioral faithfulness the missing evaluation criterion for distilled memory
- Ingest: Learning By Writing (ingest-report) - Karnofsky's writing-centered investigation loop adds the working hypothesis as control state for evidence allocation, bounded to first-person human practice
- Ingest: Lessons from Building AI Agents for Financial Services (ingest-report) - Production financial-agent report supports S3-first files with derived PostgreSQL, skill shadowing for customization, and 'model eats scaffolding,' with fiscal-period normalization as a calculator counterexample
- Ingest: Letta (MemGPT): Stateful Agents with Self-Managed Memory (ingest-report) - Agent memory platform where the LLM self-manages a three-tier memory hierarchy (core/recall/archival) using an OS analogy — the strongest existing exemplar of the agent-self-managed agency model, now evolving toward git-backed memory files
- Ingest: LibContinual: A Comprehensive Library towards Realistic Continual Learning (ingest-report) - LibContinual exposes data access, retained-state accounting, and task-stream structure as hidden privileges in continual-learning evaluation
- Ingest: LLM Knowledge Bases (ingest-report) - Karpathy on agent-maintained research wikis in Obsidian — index files and brief summaries replacing fancy RAG at roughly 100-article scale
- Ingest: LLM Position Bias Benchmark (Swapped-Order Pairwise Judging) (ingest-report) - Swapped-order pairwise-judging benchmark showing that across 27 LLM judges the median model flips its underlying winner in 44.8% of decisive cases, with large model-dependent first-position lifts on sibling story-edit pairs
- Ingest: LLM Wiki (ingest-report) - Karpathy's long-form agent-maintained wiki manifesto — explicit raw/wiki/schema architecture plus index/log separation beyond his earlier X-post workflow sketch
- Ingest: LLMs can’t jump (ingest-report) - Position paper separating deduction within supplied axioms from abductive premise invention, using Einstein's path to general relativity to motivate action-controllable world models
- Ingest: Locating and Editing Factual Associations in GPT (ingest-report) - ROME localizes factual associations to mid-layer MLPs and edits them with a rank-one weight update — the nearest approach to our named explicit-retention falsifier, but it lacks scope-addressability
- Ingest: Machine Studying (ingest-report) - Machine Studying defines corpus-only pre-task adaptation and evaluates it across inference budgets, but its preliminary interventions and fixed StudyBench decomposition support narrower claims than the headline
- Ingest: Manage Innovation Programs With a Rolling Wave (ingest-report) - Githens presents rolling-wave planning as staged uncertainty reduction for development work and a bounded example of deferring detail until information arrives.
- Ingest: Many AI analysts, one dataset (ingest-report) - Crossed 4,946-run agent-analysis experiment separates sampling dispersion from model/persona steering and finds that auditing reduces but does not remove selective-reporting risk
- Ingest: Maps (Hidden Classes) in V8 (ingest-report) - Official V8 walkthrough showing how Maps encode object layouts and how mutating a field compiled as constant invalidates dependent optimized code
- Ingest: Mathematical discoveries from program search with large language models (ingest-report) - FunSearch supplies bounded evidence for frozen-model search that selects and reuses localized programs while its evaluator and skeleton remain fixed.
- Ingest: Maximum Effective Context Window (ingest-report) - Study across 11 frontier LLMs finds maximum effective context windows up to 99% below advertised limits, task-dependent, with hallucinations approaching 100% beyond the effective window
- Ingest: MCDP 1, Warfighting (1997) (ingest-report) - Official Marine Corps doctrine grounds intent-framed delegation under uncertainty while limiting transfer to systems with competent executors, trust, and shared context.
- Ingest: Mem0: Universal Memory Layer for AI Agents (ingest-report) - Mem0's two-phase add pipeline (extract facts + LLM-judged CRUD reconciliation) is the purest production example of automated accretion-without-synthesis — now contextualized by the comparative review
- Ingest: Memento-Skills: Let Agents Design Agents (ingest-report) - Memento-Skills supplies empirical evidence for frozen-LLM deploy-time learning through routed, rewritten executable skill memory, with transfer bounded by domain alignment
- Ingest: Memory Intelligence Agent (ingest-report) - MIA mixed-substrate deep-research agent memory paper — search trajectories become both workflow memory and Planner weight updates during test-time learning
- Ingest: Memory Scaling for AI Agents (ingest-report) - Databricks memory-scaling experiments showing enterprise agent gains from external memory only when retrieval, distillation, and governance scale with the store
- Ingest: Mesa Optimizers and Language Recursion (ingest-report) - Speculative essay arguing mesa optimizers may emerge suddenly because language recursion and learned search both compress many cases into reusable generative rules.
- Ingest: Meta-Harness: End-to-End Optimization of Model Harnesses (ingest-report) - Controlled ablation showing raw execution traces (10 MTok/iter) outperform summaries by 10+ points in automated harness search — first empirical evidence for diagnostic richness as binding constraint
- Ingest: Metaobject Protocols: Why We Want Them and What Else They Can Do (ingest-report) - Primary 1993 attestation that reflective language change can be routed through a documented, explicitly marked metaobject protocol while ordinary base-level syntax and defaults remain intact
- Ingest: Minimum Viable Ontology / Domain Maps (ingest-report) - Tweet thread proposing "minimum viable ontology" — the smallest term list to orient a newcomer in a domain — with a vibecoded prototype (domainmaps.co) and pedagogical framing via "conceptual thresholds"
- Ingest: Model Discovery Agent: LLM-assisted Bayesian Experiment Design (ingest-report) - MDA splits scientific model discovery between LLM hypothesis proposal and Bayesian selection, improving interventional forecasts inside a fixed domain decomposition
- Ingest: Monkey patch (ingest-report) - Tertiary vocabulary evidence that 'monkey patch' names runtime class or module modification as a workaround and carries warnings about incompatibility, conflicts, hidden behavior, and patch warfare
- Ingest: Multi-Agent Memory from a Computer Architecture Perspective (ingest-report) - Computer-architecture analogy for multi-agent memory — shared/distributed paradigms, three-layer hierarchy, consistency protocols as the critical unsolved problem
- Ingest: Natural-Language Agent Harnesses (ingest-report) - NLAH paper externalizes agent control logic as portable natural-language artifacts — key empirical finding: explicit structure helps only when it tightens alignment with evaluator acceptance criteria, not by adding process layers
- Ingest: Nested Learning (ingest-report) - Nested Learning recasts architectures and optimizers as nested associative memories, but Hope's broad gains vary local components within a fixed weight-only decomposition
- Ingest: Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents (ingest-report) - Enterprise self-evolving-agent architecture linking decision-shaped traces to governed selection among memory, skills, harnesses, tools, and weights
- Ingest: Nick Milo's definition of Maps of Content (ingest-report) - Nick Milo defines a Map of Content as contextual mapping through clustered links, supporting a narrow navigation claim but not claimed cognitive effects.
- Ingest: Niklas Luhmann Archive keyword registers (ingest-report) - The archive documents Luhmann's keyword registers as deliberately non-exhaustive entry-point indexes, supporting a narrow historical comparison rather than general PKM claims.
- Ingest: Novel Memory Forgetting Techniques for Autonomous AI Agents (ingest-report) - Formula-based adaptive forgetting with constrained optimization for agent memory — the inspectable alternative to RL-trained memory policy, with empirical evidence that uncontrolled accumulation causes false memory propagation
- Ingest: On Doctors, Mechanics and Computer Specialists — Where are the Problems with Credence Goods? (ingest-report) - Credence-goods microeconomics paper — evidence and refinement for the verification-boundary claim, adding liability as an institutional substitute for verifiability
- Ingest: On Learning How to Learn Learning Strategies (ingest-report) - Schmidhuber's reward-gated self-modification report as historical evidence for oracle-dependent behavior learning and reversible promotion
- Ingest: On the "Induction Bias" in Sequence Models (ingest-report) - 190k-run empirical study showing transformers need orders-of-magnitude more data than RNNs for state tracking due to absence of step-by-step induction bias; introduces sharing factor kappa quantifying cross-length mechanism reuse
- Ingest: On the Criteria To Be Used in Decomposing Systems into Modules (ingest-report) - Parnas's KWIC comparison anchors why module boundaries should hide likely changes rather than mirror processing steps, while leaving downstream validation costs unproven.
- Ingest: On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification (ingest-report) - Repeated-run and shuffled-order evidence that textual-memory agents amplify variance and can turn a hidden curriculum into apparent self-improvement
- Ingest: Online Convex Programming and Generalized Infinitesimal Gradient Ascent (ingest-report) - Zinkevich's online convex-programming regret bounds provide a technical basis for direct, gateless evidence-responsive updates and their assumptions.
- Ingest: Open continual learning that you fully control (ingest-report) - A short X design sketch that turns organizational expertise into evals, environments, agent behavior, and trace-driven continual improvement
- Ingest: OpenClaw-RL: Train Any Agent Simply by Talking (ingest-report) - Deployment-time RL uses user replies, tool output, terminal feedback, and GUI state as next-state signals, collapsing training and deployment; the KB's retained-artifacts note now cites it as a deployment-time weight-update counterexample
- Ingest: Orchestrate subagents at scale with dynamic workflows (ingest-report) - Claude Code dynamic-workflows docs — the saveable, script-authored counterpoint to RLM's ephemeral orchestrators; evidence for the orchestration/run-state persistence cluster
- Ingest: Organizational Learning and Management Information Systems (ingest-report) - Argyris grounds single- and double-loop learning and theories-in-use, while warning that abstract control systems can make organizational error correction self-sealing.
- Ingest: Palantir Ontology vs Decision Traces (ingest-report) - Jaya Gupta frames Palantir-style top-down ontology and workflow-first decision traces as two ways to build LLM-facing world models
- Ingest: PMI Lexicon of Project Management Terms, Version 4.0 (ingest-report) - PMI defines progressive elaboration and rolling wave planning, anchoring Commonplace's information-timed planning claims while separating the core terms from broader methods.
- Ingest: Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems (ingest-report) - Position paper restates human-inclusive evaluation for scientific AI, but the captured page neither resolves contribution attribution nor exposes enough evidence to advance KB methodology
- Ingest: Post by @deepfates — LLM "memory" as context stuffing (ingest-report) - Deepfates argues LLM "memory" is just context-stuffing that creates false salience (Chekov's gun), advocates agentic context-building, but concludes weight updates are necessary — directly contradicts this KB's durability-not-weights position
- Ingest: Post by @koylanai (ingest-report) - Argues that pairwise judging plus round-robin win rates is a better evaluation primitive than absolute scoring for open-ended LLM tasks with no hard ground truth
- Ingest: Prime Agent: A Self-Improving RLM Harness (ingest-report) - Code-grounded analysis of Prime Agent as a persistent recursive harness, with its refinement governance failure and bundle-level evaluation limits
- Ingest: Professional Software Developers Don't Vibe, They Control (ingest-report) - Empirical study (N=112) finding experienced developers control AI agents through SE practices, not vibe coding -- grounds constraining, underspecification, and programming-practices-transfer arguments
- Ingest: Programming as Theory Building (ingest-report) - Naur's theory-building view makes maintainability depend on situated design understanding while bounding what retained rationale alone can transfer
- Ingest: Prompt Stability in Code LLMs (ingest-report) - Code-LLM study finds performance and prompt stability are distinct, smaller models may be steadier, and emotional/personality variants expose confidence miscalibration missed by standard benchmarks
- Ingest: PROV-Overview — An Overview of the PROV Family of Documents (ingest-report) - W3C PROV family roadmap — the canonical standard for 'full provenance' that the KB's lineage concept deliberately trims down from
- Ingest: Psychology already solved AI memory — identity isn't stored, it's constructed (ingest-report) - Thread proposing five psychology principles (Conway, Damasio, Bruner, Klein & Nichols) for AI memory as identity construction — directly engages the KB's open question about whether cognitive science analogies are decorative or mechanistic
- Ingest: Putting Ideas into Words (ingest-report) - Graham separates epistemic composition into exact-word commitment and neutral-reader rereading, but his human self-report does not establish an agent-learning mechanism
- Ingest: Reading Between the Dots: Decoding Hidden Computation across Filler Tokens (ingest-report) - Filler-token study as evidence that monitorability depends on the observation surface and that hidden process traces can diagnose failures
- Ingest: Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills (ingest-report) - Passive distillation of existing agent trajectories into domain skills recovers much of a reasoning mode's performance at lower output-token cost, with skill and decomposition variance untested
- Ingest: Recursive Experiential–Working Memory Evolution (ingest-report) - Recuris couples verified working state, event-triggered skill retrieval, and gated component patches; code confirms the mechanisms, while benchmark gains remain paper-only.
- Ingest: Recursive Language Models - what finally gave me the 'aha' moment (ingest-report) - Practitioner comparison of direct generation, RAG, ReAct, CodeAct, subagents, and RLM gives concrete evidence for REPL substrate, symbolic variable return, and scaffold-level truncation
- Ingest: Reducing rework with set-based systems engineering (ingest-report) - Set-based requirements, targeted experiments, and evidence-gated convergence are proposed as remedies for three causes of late systems-engineering rework.
- Ingest: Reflection and Semantics in Lisp (ingest-report) - Foundational procedural-reflection vocabulary: embedded self-theory, two-way causal connection, reflective vantage point, and explicit-versus-absorbed processor state
- Ingest: Release Handling (ingest-report) - Official Erlang/OTP attestation that live definition change is governed as deployment through versioned appup/relup plans, synchronized state migration, rollback, and explicit permanence
- Ingest: Researchers Asked LLMs for Strategic Advice. They Got "Trendslop" in Return. (ingest-report) - HBR trendslop article: LLM strategy advice follows fashionable management discourse despite prompt and context variation.
- Ingest: ResNet revolution (ingest-report) - Residual learning as architecture-enabled scale and a warning about benchmark targets masquerading as capabilities
- Ingest: Safe superintelligence (ingest-report) - Operational intelligence measures and safety narratives tested against objectives, proxy warrant, and oracle domains
- Ingest: Scaling Managed Agents: Decoupling the brain from the hands (ingest-report) - Anthropic Managed Agents report showing brain/hand/session interface decomposition, durable session logs, and stale harness assumptions as model capability changes
- Ingest: ScienceFlow: A Long-horizon Agent for ML Research, Scientific Discovery and Beyond (ingest-report) - ScienceFlow implements recoverable research workspaces, evidence-gated checkpoints, bounded memory, and resource control, but its benchmark gains remain unreproduced and decomposition-bound
- Ingest: Second-Order Cybernetics: An Historical Introduction (ingest-report) - Historical placement of observed versus observing systems and the observer's epistemological inclusion
- Ingest: Self-Harness: Harnesses That Improve Themselves (ingest-report) - Same-model harness optimization improves three Terminal-Bench agents through structured failure mining and regression-gated local edits, within a tightly fixed outer method
- Ingest: Self-Improving AI Coding Agents Through Accumulated Behavioral Rules (ingest-report) - Production evidence for review-derived coding-agent rules, bounded by an instance-to-rule oracle gap and an uncontrolled, fixed-architecture deployment.
- Ingest: Self-Improving Algorithms (ingest-report) - Ailon et al. learn a product input distribution, then retain data structures that reach entropy-optimal limiting performance
- Ingest: Self-Revising Discovery Systems for Science (ingest-report) - Category-theoretic discovery framework: typed-artifact copresheaves, retrieval/search/discovery typology, and MDL/AIC gates
- Ingest: Self-training Large Language Models through Knowledge Detection (ingest-report) - EMNLP paper turning unknown-detection scores into filtered DPO preference data, with selective self-training reducing hallucination and limiting forgetting on Wikipedia QA
- Ingest: SELF: The Power of Simplicity (ingest-report) - Ungar and Smith's primary classless-OO attestation: Self rejects the class/instance split for prototypes and parent-based sharing, gaining cloning and per-object flexibility while losing explicit organizational cues
- Ingest: Simplicity, hidden in complexity (ingest-report) - Compression as predictive objective, measurement instrument, and interface—plus its limits as a theory of generalization
- Ingest: Simplification (ingest-report) - Mike Briggs's practitioner account makes vague simplification requests evidence for underspecification and treats agent architectural hunches as candidacy, not verdict, evidence
- Ingest: Skill Synthesis — Materializing Knowledge as Skills (ingest-report) - Sentry co-founder's practitioner report on synthesizing Claude Code skills from domain-specific source material (commit history, security patches, OWASP docs) — found 8 real IDORs missed by professional pen testing
- Ingest: SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning (ingest-report) - SkillRL paper showing trajectory-distilled skill banks co-evolving with GRPO-trained agent policy, bridging readable skill artifacts and weight-based learning
- Ingest: Slate: Moving Beyond ReAct and RLM (ingest-report) - Practitioner architecture uses bounded worker threads that return compressed episodes to an orchestrator, converging on Commonplace's bounded-context model for memory, coherence, and decomposition
- Ingest: Software Engineering of Self-Adaptive Systems: An Organised Tour and Future Challenges (ingest-report) - Weyns's six-wave engineering perspective supplies a bounded self-adaptation vocabulary and a warning that MAPE-K is an engineering reference model, not a membership definition.
- Ingest: Solving a Million-Step LLM Task with Zero Errors (ingest-report) - MAKER achieves zero errors over one million LLM steps via maximal decomposition into single-step microagents with first-to-ahead-by-k voting and red-flagging — proves O(s ln s) cost scaling when hard per-step oracles exist
- Ingest: Spacebot: AI Agent for Teams and Communities (ingest-report) - Spacebot README ingest covering process-typed concurrent agent runtime architecture, branch scoping, cortex supervision, and typed unified memory
- Ingest: SPADE: Self-Play in Adaptive Synthetic Executable Environments (ingest-report) - SPADE is a code-grounded case of adaptive executable-environment generation inside a fixed training decomposition, with released mechanisms but unreleased outcome artifacts at the pinned commit.
- Ingest: Structured Test-Time Scaling: From Multi-Agent Systems to General Inference Architectures (ingest-report) - Formal result connects topology compression, scope isolation, and verification as a causal chain enabling hierarchical multi-agent systems to avoid exponential error accumulation
- Ingest: SuperARC — Can Complexity and Uncomputability Explain Intelligence? (ingest-report) - Ingest of SuperARC — AIT-grounded benchmark where frontier LLMs score phi ~0.03 while neuro-symbolic CTM/BDM achieves 1.000 on recursive compression; newer models regress; print-statement-only outputs demonstrate zero algorithmic abstraction
- Ingest: Sutton and Javed on why AI models stop learning (ingest-report) - Sutton and Javed argue that context-state adaptation cannot replace continual weight learning, posing a counterpoint to Commonplace's whole-system account.
- Ingest: SWE-bench Science full paper on scientific-software repairs (ingest-report) - Author-controlled benchmark paper separates visible test success from scientific correctness and bounds the mixed effects of supplied scientific guidance.
- Ingest: SWE-bench Science on coding-agent repairs (ingest-report) - Scientific-software benchmark separates public-test conformance from private scientific correctness and finds supplied guidance can help, fail, or anchor repairs.
- Ingest: Symbiotic Reflection between an Object-Oriented and a Logic Programming Language (ingest-report) - Primary precursor defining symbiotic reflection, mutual cross-language reasoning and action, and upping/downing entity transfer.
- Ingest: Symbolic Learning Enables Self-Evolving Agents (ingest-report) - Early whole-harness optimizer treats prompts, tools, and pipeline topology as jointly learnable language-mediated artifacts, with cross-node credit assignment and same-oracle rollback
- Ingest: The "Mismanaged Geniuses" Hypothesis (ingest-report) - Hypothesis that current frontier LMs are bottlenecked by learned decomposition/scaffold policy rather than base capability, using RLMs and orchestrator-subagent systems as evidence
- Ingest: The Agent Loop Architecture (ingest-report) - Inngest practitioner framing of durable agent loops as loop + skill + orchestrator, useful for the run-state and skill-persistence boundary
- Ingest: The Alberta Plan for AI Research (ingest-report) - A continual-learning agent architecture and research roadmap that separates per-step adaptation, model-based planning, and progressively learned abstractions.
- Ingest: The AlexNet moment (ingest-report) - AlexNet as a worked case of moving representation choices inside the effective update space
- Ingest: The Anatomy of a Design Theory (ingest-report) - Gregor and Jones's six-core/two-additional anatomy separates design-theory content from implementation agents and physical instantiations
- Ingest: The Anatomy of an Agent Harness (ingest-report) - Practitioner taxonomy deriving harness components (filesystem, bash, sandboxes, memory, context management, long-horizon execution) from model limitations — provides the component anatomy that bridges Lopopolo's practice and the cybernetics framing
- Ingest: The Best Model Routing is Task Specific (ingest-report) - Jerry Liu on task-specific model routing: generic gateways handle provider routing, but workflow-specific routers earn cost and accuracy gains from private evals and input taxonomies
- Ingest: The birth of hyperscale (ingest-report) - Scaling laws as regime-conditional engineering guidance rather than unconditional capability laws
- Ingest: The Bitter Lesson (ingest-report) - Sutton argues that scalable search and learning outlast hand-built domain knowledge; this is the primary anchor for three KB interpretations.
- Ingest: The Bitter Lesson (ingest-report) - Wikipedia-contextualized capture of Sutton's Bitter Lesson, useful for scaling arguments and caveats about general methods versus hand-coded knowledge.
- Ingest: The Bug That Shipped (ingest-report) - 3,700-trial practitioner evidence that coding models can diagnose deployment failures when explicitly probed but rarely surface them in undirected self-review
- Ingest: The File System Is the New Database: How I Built a Personal OS for AI Agents (ingest-report) - Practitioner report on a file-based personal OS for AI agents, useful as self-reported evidence for filesystem-first context engineering.
- Ingest: The Flawed Ephemeral Software Hypothesis (ingest-report) - Essay distinguishing vibe coding from true software ephemerality, arguing that state, integration, interface stability, and auditability keep important systems anchored to durable artifact stacks.
- Ingest: The Geometry of Forgetting (ingest-report) - Embedding-memory paper arguing that interference and low effective dimensionality, not time decay, drive forgetting and false recall in similarity retrieval.
- Ingest: The GitHub for Context Doesn't Exist Yet (ingest-report) - Prukalpa on shared organizational context outliving agent-stack churn, with production failures motivating lineage, semantic impact review, controlled trace learning, security, and portability
- Ingest: The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations (ingest-report) - Netflix case study of rationale-aware rubric tuning, dual-role judge deployment, and human-calibrated drift monitoring for recommendation explanations
- Ingest: The Log Is the Agent (ingest-report) - Omnara essay arguing an agent IS its durable append-only event log; log ownership = agent ownership
- Ingest: The Nature of Theory in Information Systems (ingest-report) - Gregor's mature taxonomy defines design-and-action theory by prescriptive purpose and exposes actionability as a separate theory-actor-context relation
- Ingest: The No Free Lunch Theorem (ingest-report) - No Free Lunch explainer grounding codification as an unavoidable inductive-bias bet: strategies win only when their assumptions match the problem distribution.
- Ingest: The Optimal Choice of Hypothesis Is the Weakest, Not the Shortest (ingest-report) - Bennett's extension-size alternative to minimum description length, bounded by its uniform task prior and fixed representational decomposition
- Ingest: The Perfect Search Engine Is Not Enough (ingest-report) - A CHI diary study gives bounded human evidence that contextual local steps complement direct keyword jumps, without testing LLM agents or KB metadata designs.
- Ingest: The Physics of Symbols: Bridging the Epistemic Cut (ingest-report) - Epistemic-cut vocabulary for symbolic constraints, material dynamics, and the limits of natural-language/code analogies
- Ingest: The pivot to reasoning (ingest-report) - Reasoning architectures and runtime methods read through fixed decompositions, process validity, and benchmark turnover
- Ingest: The Price of Meaning: Why Every Semantic Memory System Forgets (ingest-report) - Formal no-escape theorem for semantic memory interference, with exact-record and symbolic-verifier escape clauses that sharpen retrieval-vs-verification tradeoffs.
- Ingest: The Risks of Invariant Risk Minimization (ICLR 2021 camera-ready) (ingest-report) - Full ICLR 2021 text behind the KB's IRM counterweight: the E > d_e threshold with a matching lower bound, the non-linear construction that reverts to ERM at test, and the critique of Arjovsky et al.'s bound
- Ingest: The Second Brain Trap (ingest-report) - PlugLab AI founder reframes "second brain" failure as stored knowledge that never activates in context, then proposes trigger-rich graph structure as the fix
- Ingest: The Self-Healing Agent Harness (ingest-report) - Practitioner report where live agent-response graders feed tickets, auto-fixes, re-grading, and rollout gates instead of sitting as offline eval dashboards
- Ingest: The Spec Is the New Code — A Guide to Spec Driven Development (ingest-report) - MercadoLibre engineering lead's practitioner guide to Spec Driven Development — the spec/plan/task/implement cascade as methodology for eliminating agent ambiguity, with ecosystem convergence evidence and maturity-level progression
- Ingest: The Use of Proximal Information Scent to Forage for Distal Content on the World Wide Web (ingest-report) - Information scent as a Brunswikian lens model — judging unseen content from proximal cues, with spreading activation and a random-utility rule fitted to Web protocols; grounds the KB's pointer-quality and stopping claims
- Ingest: The Vision of Autonomic Computing (ingest-report) - Kephart & Chess 2003, origin of MAPE-K and the four self-* properties; the primary source behind the KB's reference-model-not-definition reading of self-adaptive loops.
- Ingest: The What & When of Self-Evolving Agents (ingest-report) - 3×3 framework mapping self-evolving-agent updates across external files, harnesses, and model weights over task, session, and population horizons
- Ingest: The Y-Combinator for LLMs (ingest-report) - λ-RLM preprint replacing open-ended RLM REPL code with typed combinators, formal bounds, and long-context benchmark evidence
- Ingest: Thread by @frgx (ingest-report) - Figma's practitioner account of turning security precedents into a trusted, tested policy artifact that drives agent review, repo auditing, and secure code generation.
- Ingest: Thread by @LechMazur (ingest-report) - Lech Mazur's public benchmark announcement compressing the headline position-bias result, the GPT-5.4 callout, and the operational motivation from everyday comparison prompts
- Ingest: Three Dimensions That Matter To An Agent Memory Store (ingest-report) - Opinionated agent-memory design essay: hybrid retrieval, selective push, and organization-scoped storage, with anti-graph claims that outrun its evidence
- Ingest: ToolGate: Contract-Grounded and Verified Tool Execution for LLMs (ingest-report) - ToolGate makes tool-state mutation transactional through typed symbolic state, pre-execution admissibility checks, and post-execution result verification
- Ingest: Toulmin Argument (ingest-report) - Pedagogical treatment of Toulmin's six-part argument model — canonical source for the structured-claim type's Evidence/Reasoning/Caveats sections
- Ingest: Toward measuring recursive self-improvement (ingest-report) - Shōbench's whole-state before/after study detects five clean gains but measures retained self-directed adaptation rather than recursive compounding
- Ingest: Towards a Science of AI Agent Reliability (ingest-report) - Reliability framework paper arguing mean task success is inadequate for agents, replacing it with consistency, robustness, predictability, and safety.
- Ingest: Towards a Science of Scaling Agent Systems (ingest-report) - Controlled multi-agent scaling paper showing coordination gains depend on task decomposability, verification, and context overhead rather than agent count.
- Ingest: Towards Automating Eval Engineering (ingest-report) - Eval Engineering Skill announcement describing trace mining, user-guided task design, Harbor environments, verifier inspection, and iterative agent improvement
- Ingest: Towards Automating Scientific Review with Google's Paper Assistant Tool (ingest-report) - Google PAT paper as evidence for verifiable-subrole review automation: segmenting manuscripts, scaling inference, and keeping humans accountable for final review authority.
- Ingest: Towards Causal Representation Learning (ingest-report) - Causal representation learning grounds the claim that causal models support intervention, counterfactual, and reusable-mechanism generalization
- Ingest: Towards Faithfully Interpretable NLP Systems (ingest-report) - Jacovi and Goldberg separate faithfulness from plausibility, expose the assumptions behind common explanation tests, and argue for graded rather than binary faithfulness
- Ingest: Towards Linguistic Symbiosis of an Object-Oriented and a Logic Programming Language (ingest-report) - Implementation evidence that linguistic symbiosis across logic and object-oriented paradigms requires explicit syntactic and semantic mappings.
- Ingest: TRACE: TRajectory Attribution for Automated Context Engineering (ingest-report) - TRACE supports staged, source-verified context diagnosis but measures recommendation accuracy rather than applied repair
- Ingest: tracecraft (ingest-report) - S3-backed CLI coordination tool for multi-agent systems — exposes how file-backed coordination depends on workload-specific consistency and authority requirements
- Ingest: Trajectory-Informed Memory Generation for Self-Improving Agent Systems (ingest-report) - IBM pipeline extracts strategy, recovery, and optimization tips from trajectories for runtime retrieval; subtask granularity yields +14.3-point AppWorld gains under a narrow task-completion oracle
- Ingest: Transformers Learn In-Context by Gradient Descent (ingest-report) - Mechanistic ICML paper showing in-context regression can be implemented as gradient descent inside Transformer forward passes, sharpening the internal half of the KB's in-context-learning theory
- Ingest: Verbalizable Representations Form a Global Workspace in Language Models (ingest-report) - Anthropic J-space paper as evidence for probeable parametric state, activation-vs-presence, and externalized reasoning as internal-workspace relief
- Ingest: Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures (ingest-report) - Controlled simulation finds that effect verification before retry cuts duplicate tool actions, while a fixed simulator and bundled treatment limit attribution
- Ingest: We Should Take Text Optimization More Seriously (ingest-report) - Manifesto arguing text optimization is a legitimate, sample-efficient update mechanism enabling update-time compute — upstream restatement of the KB continual-learning cluster
- Ingest: What did Ilya see? (ingest-report) - A selective AI canon as evidence about worldview, orientation, and omitted design alternatives
- Ingest: What is an Agent Harness (ingest-report) - Arize taxonomy separates nine harness components and adds permission/safety plus lifecycle hooks as first-class primitives, a fifth independent convergence on the harness decomposition
- Ingest: What spec-driven development gets wrong (ingest-report) - Augment's argument that spec-driven development fails unless agents co-maintain the spec — bidirectional spec as a mechanism for matching maintenance throughput to generation throughput
- Ingest: What Survives in Multi-Agent Systems (ingest-report) - Applied bitter-lesson analysis predicting which multi-agent patterns survive stronger models — argues filesystem, forking, and spawning are structural while fixed orchestration is a vision feature
- Ingest: When code is free, research is all that matters (ingest-report) - Investor/researcher argument that oracle availability (not capability) determines automation boundary for cognitive work — research taste is unautomatable because problem selection has no ground truth
- Ingest: When is it better to think without words? (ingest-report) - Karlsson decomposes human thinking into wordless exploration and written testing, stabilization, and relay, while leaving the neuroscience and LLM analogy speculative
- Ingest: Where It Lives Is Not What It Is (architectural vocabulary for retained adaptation) (ingest-report) - External paper derived from Commonplace's four-field artifact analysis; the ingest tracks paper-from-notes lineage, adds sovereignty-risk refinements, and may become a citable authority after acceptance
- Ingest: Where It Lives Is Not What It Is (June 2026 version) (ingest-report) - Updated self-authored ASISAS position paper adding 141-system corpus evidence to the four-field retained-artifact vocabulary.
- Ingest: Why AI systems don't learn and what to do about it (ingest-report) - Position paper arguing current AI externalizes learning into human-run MLOps and proposing an A-B-M architecture where meta-control arbitrates observation and action learning for lifelong adaptation.
- Ingest: Why Large Language Models Fail at Tabular Prediction (ingest-report) - Pure-inference tabular experiments separate column readability from multi-column integration and provide a reusable benchmark-contamination probe
- Ingest: Why LLMs can’t make your code simpler (ingest-report) - Answer.AI's complexity argument is a bounded case of maintainability receiving weaker selection pressure than correctness, with ADRs able to supply the missing design context
- Ingest: Why Multi-Agent Pipelines Fail for Complex Analytics (ingest-report) - A practitioner redesign centralizes diagnostic authority while retaining fact-returning sub-agents, deterministic anomaly detection, and graph-bounded hypothesis traversal
- Ingest: Why Software Factories Fail (ingest-report) - Dex Horthy's failed lights-off software-factory case supports the verification boundary and challenges agent-only maintainability review
- Ingest: Why Software Factories Fail: Benchmarking the new frontier (ingest-report) - Dex Horthy's SlopCodeBench run turns long-horizon code maintenance into a delayed benchmark signal while showing why deterministic proxies do not yet warrant lights-off autonomy
- Ingest: Why Software Factories Fail: Turning the lights back on (ingest-report) - Dex Horthy's staged human-review workflow moves maintainability judgment upstream and limits unchecked agent work with vertical slices
- Ingest: Why You Should Almost Never Use AI to Write Anything Substantive (ingest-report) - Grunewald's critique exposes passive-assent risk and a gap between expert interpretation of a writing commission and generic LLM completion
- Ingest: Workspace Optimization: How to Train Your Agent (ingest-report) - DreamTeam adapts a fixed-model agent by revising typed code and role context from prediction failures, but does not isolate the value of its fixed workspace decomposition.
- Ingest: Your Old Agent Architecture Is Dead… Meet Its Replacement (ingest-report) - A strong architectural argument for recursive code–inference execution, plan-as-program workflows, and compiling natural-language guardrails into typed deterministic middleware.
- Writing conventions for kb/sources/