Ingest: Gentle-Coding Comparative Research Catalog

Type: kb/sources/types/ingest-report.md

Classification

This repository document is an annotated bibliography organized as project research infrastructure, not an empirical study or literature review with a stated search method. Author: the Gentle-Coding project maintainers provide topical curation and brief interpretations, but the document does not establish their research credentials or independent review process.

Summary

The catalog groups links about emotional prompting, model stress, sycophancy, and human interaction into sub-network, inference, and multi-agent layers. It is most useful as a discovery map for selecting primary sources and designing comparisons across interaction and evaluation boundaries. Readers should not rely on its confident one-line summaries as evidence: the document reports no search protocol, source-verification procedure, experiments, or synthesis method.

Quotes

No source quotes have been retained yet.

Connections Found

The source is a lead index and methodological limitation, not an evidence anchor. Its agreement-pressure and long-context entries compare with context contamination operates below an agent's compliance reasoning, but they do not establish a shared mechanism. Its warnings about missing human perception and potentially biased evaluator models compare with evaluation automation is phase-gated by comprehension. Language Models, Like Humans, Show Content Effects on Reasoning Tasks is the grounded methodological counterpoint because it measures human/model convergences and divergences that this catalog only groups conceptually.

Extractable Value

  1. A three-layer retrieval frame for pressure-related model behavior -- The sub-network, inference, and multi-agent split can help route primary-source investigations by where a claimed effect enters an agent system, without treating the split itself as validated. [just-a-reference]
  2. A focused primary-source queue for sycophancy evaluation -- The human-in-the-loop, evaluator-bias, agreement-pressure, and context-length entries identify concrete studies and benchmarks to capture independently before the KB makes claims about stance drift. [deep-dive]
  3. A distinction to test between social agreement pressure and general context contamination -- The catalog supplies candidate social-pressure interventions that could be compared with the controlled stance-influence behavior already represented in the KB. [experiment]
  4. An evaluation-design warning -- Its multi-agent grouping foregrounds the possibility that automated judges reproduce the behavior being measured, making human perception and judge calibration explicit checks for any later sycophancy assay. [quick-win]
  5. A human/model comparison checklist -- The catalog's paired framing is useful for asking where behavioral parallels diverge, while the existing content-effects ingest supplies the stronger example of measuring rather than assuming those parallels. [just-a-reference]

Limitations (our opinion)

The document is a point-in-time repository-maintainer bibliography whose entries compress heterogeneous papers, blogs, repositories, and project claims into assertive one-line conclusions. It provides no inclusion criteria, verification record, quotations, methods comparison, or treatment of conflicting findings; several entries are future-dated relative to capture, and malformed table rows further weaken confidence in curation. The human/LLM layer analogy may also conflate similar language with a shared mechanism. Consequently, the catalog can guide discovery and experiment design but cannot support promotion of its summaries as source-grounded claims.

Independently capture and ingest the linked paper “Sycophancy Claims About Language Models: The Missing Human-in-the-Loop” to verify its evaluation taxonomy and human-perception critique against the primary text.