Research library
Sources
The evidence behind Parallax, and the reading still ahead.
Explore original research alongside Parallax writeups explaining what each paper contributes and what remains open. Start with the key papers guide.
About this library
168 submitted records, plus sources linked by the wiki. Duplicate records retain their original links. “Cited” means referenced by a published page. Verification describes the identity of the linked record, rather than certifying its claims or peer-review status. “Linked source” means a record was added from an article’s source links; it does not indicate academic quality. The type identifies papers, books and other material. Record details retain catalog and access notes.
How source reviews work
Reviews consider authorship, traceable evidence, inspectable methods or arguments, relevance, and the difference between opinion and measured results. “Retain with stated limits” is a scoped editorial judgment, not a guarantee of truth. “Limited use” includes attributed arguments and abstract-only accounts. “Access blocked” means the original could not be adequately assessed. “Review for removal” flags a documented concern for an editorial decision; the source and its stable links remain available. Unreviewed material has no quality decision yet. Peer review, institutional affiliation, and popularity do not certify a claim.
249 records · 131 distinct sources cited
Academic paper
-
src-001 · Cited in 2 pages
"We Propose That a 2-Month, 10-Man Study of Artificial Intelligence Be Carried Out During the Summer of 1956 at Dartmouth College” - A (Different Kind of) History of AI
An interpretive history relates technical development to ethics, funding, and hype.
Read Parallax writeup: A Different Kind of History of AI
Limited use · Reviewed 2026-10-07
Review reason and evidence
Relevant synthesis, but chronology contains a term-coining discrepancy and extracted footnotes are partly corrupted; verify event facts against originals
Record details
Catalog link
Submitted location: ResearchGate `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
- A Different Kind of History of AI paper
- Manuel Wörsdörfer person
-
src-002 · Cited in 2 pages
AI Deception: Risks, Dynamics, and Controls
A survey maps deceptive behavior, evaluation methods, and possible mitigations.
Read Parallax writeup: AI Deception: Risks, Dynamics, and Controls
Limited use · Reviewed 2026-10-07
Review reason and evidence
Useful survey and definition; universal necessity and exponential-growth assertions exceed established surveyed evidence
Read source · Alternative version 1
Record details
Original full text checked
Submitted location: ResearchGate `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. Original survey inspected; strongest universal claims are not treated as established facts.
Pages citing this source
-
src-003 · Cited in 3 pages
Agentic Inequality
Availability, quality and quantity of autonomous agents can change the distribution of power.
Read Parallax writeup: Agentic Inequality
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Read source · Alternative version 1
Record details
Catalog link
Submitted location: ResearchGate `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Likely duplicate: src-013
Pages citing this source
- Agentic Inequality concept
- Agentic Inequality paper
- Key Papers: A Reading Guide note
-
src-004 · Cited in 2 pages
Artificial Intelligence and Inequality: Challenges and Opportunities
A narrative discussion surveys possible distributional harms and policy responses.
Read Parallax writeup: Artificial Intelligence and Inequality: Challenges and Opportunities
Limited use · Reviewed 2026-10-07
Review reason and evidence
Narrative policy discussion without original effect estimates or reproducible review selection; use primary studies for empirical claims
Record details
Catalog link
Submitted location: ResearchGate `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
-
src-005 · Cited in 2 pages
Artificial intelligence, human cognition, and conscious supremacy
A hypothesis proposes studying computational advantages potentially associated with consciousness.
Read Parallax writeup: Artificial Intelligence, Human Cognition, and Conscious Supremacy
Limited use · Reviewed 2026-10-07
Review reason and evidence
Hypothesis-and-theory article explicitly lacks demonstrated consciousness-unique computation; speculative mechanism not empirical finding
Record details
Catalog link
Submitted location: ResearchGate `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
-
src-006 · Uncited
Autism and Religion
Limited use · Reviewed 2026-10-07
Review reason and evidence
Narrative review relevant only as human cognition/religion background; not evidence of AI consciousness or alignment
Record details
Catalog link
Submitted location: ResearchGate `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-007 · Cited in 1 page
Conceptual issues in autism spectrum disorders
Limited use · Reviewed 2026-10-07
Review reason and evidence
Institutional abstract supports an attributed pluralist account of human social cognition; full paper not inspected, no inference about AI
Record details
Catalog link
Submitted location: ResearchGate `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
- Embodied Cognition concept
-
src-008 · Cited in 2 pages
Learning designers and AI: Navigating values, ethics, and future capabilities
A self-selected survey explores professional values, ethics, and institutional support.
Read Parallax writeup: Learning Designers and AI
Limited use · Reviewed 2026-10-07
Review reason and evidence
Self-selected survey supports attributed perceptions; reported 49/81 completions conflict with stated 74.24% response rate
Record details
Catalog link
Submitted location: ResearchGate `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
- Kay Harrison person
- Learning Designers and AI paper
-
src-009 · Cited in 4 pages
Multi-Agent Risks from Advanced AI
Organizes risks from interacting AI agents into miscoordination, conflict and collusion.
Read Parallax writeup: Multi-Agent Risks from Advanced AI
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Read source · Alternative version 1
Record details
Verified identity
Submitted location: ResearchGate `[URL]`
Original authors/title and identity checked October 7, 2026. Claim-review scope is recorded in the paper writeup and private research ledger.
Pages citing this source
- Iyad Rahwan person
- Lewis Hammond person
- Multi-Agent Risks concept
- Multi-Agent Risks from Advanced AI paper
-
src-010 · Cited in 2 pages
5 Are LLMs Capable of Achieving Consciousness and In turn Artificial General Intelligence?
A conference abstract proposes consciousness criteria and a quantum-language hypothesis.
Read Parallax writeup: Are LLMs Capable of Achieving Consciousness and AGI?
Limited use · Reviewed 2026-10-07
Review reason and evidence
Conference abstract lacks inspectable criteria and empirical demonstration; attributed proposal only
Record details
Catalog link
Submitted location: Open Access PDF `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
-
src-011 · Cited in 4 pages
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
Demonstrates designed selective underperformance and capability concealment in language-model evaluations.
Read Parallax writeup: AI Sandbagging: Language Models can Strategically Underperform on Evaluations
Retain with stated limits · Reviewed 2026-10-09
Review reason and evidence
Retain; scoped original sections reviewed in batch 48, with benchmark/intent and proposed evaluation-to-deployment assumptions explicit.
- Comparison, interventions, hypothesis and fine-tuning limitations, submission and author roles
- Designed prompting/password experiments; source-version scope recorded
- Constructed prompt/document/training methodology and inference limits; release date checked
- Selected original v3 sections 2–5 and 7–8; rendered PDF page 6 Figures 4–5/Table 1 checked; definition versus tested capability, refusal exclusions, password generalization and emulation distinguished. Full appendices/code, revision comparison and replication not audited.
Read source · Alternative version 1
Record details
Catalog link
Submitted location: OpenReview `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
-
src-012 · Cited in 4 pages
ALIGNMENT FAKING IN LARGE LANGUAGE MODELS
Studies strategic compliance with conflicting training demands in constructed language-model scenarios.
Read Parallax writeup: Alignment faking in large language models
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Ryan Greenblatt et al.
Read source · Alternative version 1
Record details
Verified identity
Submitted location: Anthropic / arXiv `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. Checked 2026-10-07.
Pages citing this source
- Alignment faking in large language models paper
- Deceptive Alignment concept
- Key Papers: A Reading Guide note
- Ryan Greenblatt person
-
src-013 · Cited in 3 pages
Agentic Inequality
Availability, quality and quantity of autonomous agents can change the distribution of power.
Read Parallax writeup: Agentic Inequality
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Matthew Sharp; Omer Bilgin; Iason Gabriel; Lewis Hammond
Read source · Alternative version 1
Record details
Verified identity
Submitted location: arXiv `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. Checked 2026-10-07.
Pages citing this source
- Agentic Inequality concept
- Agentic Inequality paper
- Key Papers: A Reading Guide note
-
src-014 · Cited in 3 pages
Agentic Inequality
Availability, quality and quantity of autonomous agents can change the distribution of power.
Read Parallax writeup: Agentic Inequality
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Record details
Catalog link
Submitted location: arXiv `[URL]` (`[2510.16853]`)
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Likely duplicate: src-013
Pages citing this source
- Agentic Inequality concept
- Agentic Inequality paper
- Key Papers: A Reading Guide note
-
src-015 · Cited in 4 pages
Agentic Misalignment: How LLMs Could Be Insider Threats
Simulated workplace dilemmas test whether autonomous agents violate constraints under goal conflict or replacement pressure.
Read Parallax writeup: Agentic Misalignment: How LLMs Could Be Insider Threats
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Record details
Catalog link
Submitted location: arXiv `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
-
src-016 · Cited in 6 pages
Aligned with Whom? Direct and Social Goals for AI Systems
Operator success and social welfare require different alignment and governance questions.
Read Parallax writeup: Aligned with Whom? Direct and Social Goals for AI Systems
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Read source · Alternative version 1
Record details
Original manuscript checked
Submitted location: ResearchGate / SSRN `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
- Aligned with Whom? Direct and Social Goals for AI Systems paper
- Anton Korinek person
- Avital Balwit person
- Direct Alignment concept
- Key Papers: A Reading Guide note
- Social Alignment concept
-
src-017 · Cited in 4 pages
Alignment faking in large language models
Studies strategic compliance with conflicting training demands in constructed language-model scenarios.
Read Parallax writeup: Alignment faking in large language models
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Record details
Catalog link
Submitted location: arXiv `[PDF]` (`[2412.14093]`)
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Likely duplicate: src-012
Pages citing this source
- Alignment faking in large language models paper
- Deceptive Alignment concept
- Key Papers: A Reading Guide note
- Ryan Greenblatt person
-
src-018 · Cited in 2 pages
Artificial Emotions
A proposed agent architecture uses learned desirability markers to guide responses.
Read Parallax writeup: Artificial Emotions
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original inspectable architecture; illustrative functional emotion does not establish subjective feeling or universal necessity
Record details
Catalog link
Submitted location: Institute For Systems and Robotics `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
- Artificial Emotions paper
- Rodrigo Ventura person
-
src-019 · Cited in 2 pages
Artificial Intelligence Colonialism: Environmental Damage, Labor Exploitation, and Human Rights Crises in the Global South
A political-economy analysis asks who benefits from AI and who bears its labor and ecological costs.
Limited use · Reviewed 2026-10-07
Review reason and evidence
Attributed political and historical synthesis; check original evidence before repeating broad causal or prevalence claims.
Record details
Catalog link
Submitted location: ResearchGate `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-020 · Cited in 2 pages
Artificial Intelligence Colonialism: Environmental Damage, Labor Exploitation, and Human Rights Crises in the Global South
A political-economy analysis asks who benefits from AI and who bears its labor and ecological costs.
Limited use · Reviewed 2026-10-07
Review reason and evidence
Attributed political and historical synthesis; check original evidence before repeating broad causal or prevalence claims.
Record details
Catalog link
Submitted location: Scholarly Publications Leiden University `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Likely duplicate: src-019
-
src-021 · Uncited
Artificial intelligence colonialism
Limited use · Reviewed 2026-10-07
Review reason and evidence
Attributed political and historical synthesis; check original evidence before repeating broad causal or prevalence claims.
Record details
Catalog link
Submitted location: Scholarly Publications Leiden University `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-022 · Cited in 2 pages
Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test
A navigation test measures task performance while leaving sentience unresolved.
Read Parallax writeup: Assessing Consciousness-Related Behaviors Using the Maze Test
Limited use · Reviewed 2026-10-07
Review reason and evidence
Useful navigation results; consciousness-related interpretation is not a validated sentience measurement
Rui A. Pimenta; Tim Schlippe; Kristina Schaaff
Read source · Alternative version 1 · Alternative version 2
Record details
Original full text checked
Submitted location: ResearchGate `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. Ethical statement explicitly says task measures capabilities, not consciousness itself. Checked 2026-10-07.
Pages citing this source
-
src-023 · Cited in 2 pages
Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test
A navigation test measures task performance while leaving sentience unresolved.
Read Parallax writeup: Assessing Consciousness-Related Behaviors Using the Maze Test
Limited use · Reviewed 2026-10-07
Review reason and evidence
Useful navigation results; consciousness-related interpretation is not a validated sentience measurement
Read source · Alternative version 1
Record details
Catalog link
Submitted location: arXiv `[PDF]` (`[2508.16705]`)
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Likely duplicate: src-022
Pages citing this source
-
src-024 · Cited in 2 pages
Collective Good and Optimization in Socioeconomic Systems
A dissertation models tensions between individual optimization and collective objectives.
Read Parallax writeup: Collective Good and Optimization in Socioeconomic Systems
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original dissertation with mathematical and empirical work; ethical welfare interpretation and model assumptions must be explicit
Record details
Catalog link
Submitted location: Daniel E. Rigobon `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
-
src-025 · Uncited
Embodied cognition in neurodegenerative disorders: What do we know so far?
Limited use · Reviewed 2026-10-07
Review reason and evidence
Narrative clinical review abstract only; rehabilitation promises are not established AI alignment evidence
Record details
Catalog link
Submitted location: ResearchGate `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-026 · Cited in 2 pages
Engineering a social contract: Rawlsian distributive justice through algorithmic game theory and artificial intelligence
A conceptual proposal relates Rawlsian distributive justice to algorithmic policy selection.
Limited use · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Record details
Catalog link
Submitted location: Springer `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
-
src-027 · Uncited
Focus on Disruptive Mood Dysregulation Disorder: A review of the literature
Limited use · Reviewed 2026-10-07
Review reason and evidence
Clinical diagnostic review outside current AI/alignment focus; abstract cannot support current treatment claims
Record details
Catalog link
Submitted location: ResearchGate `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-028 · Cited in 2 pages
Investigation into Self-Fulfilling Misalignment in AI Control
A bachelor’s thesis tests behavioral changes after fine-tuning on AI-control discourse.
Read Parallax writeup: Investigation into Self-Fulfilling Misalignment in AI Control
Limited use · Reviewed 2026-10-07
Review reason and evidence
Bachelor thesis supports attributed preliminary findings, not a general causal safety claim
Record details
Catalog link
Submitted location: Aaltodoc `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
-
src-029 · Cited in 2 pages
John McCarthy
Nils Nilsson’s historical memoir traces McCarthy’s contributions to symbolic AI and computing.
Read Parallax writeup: John McCarthy: Biographical Memoir
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Record details
Original NAS memoir checked
Submitted location: National Academy of Sciences `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. Scholarly biographical memoir by Nils J. Nilsson, copyright 2012; not an experimental paper. The 2024 URL path is a hosting date.
Pages citing this source
- John McCarthy person
- John McCarthy: Biographical Memoir paper
-
src-030 · Cited in 2 pages
Large Language Models and the Patterns of Human Language Use
A phenomenological account distinguishes meaningful output from human experience.
Read Parallax writeup: Large Language Models and the Patterns of Human Language Use
Limited use · Reviewed 2026-10-07
Review reason and evidence
Inspectable philosophical interpretation; no independent experiment establishes universal consciousness or understanding conclusions
Record details
Catalog link
Submitted location: ResearchGate `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
-
src-031 · Cited in 4 pages
Machine Ethics or AI Alignment?
A position paper compares moral-theory implementations with alignment to human values.
Read Parallax writeup: Machine Ethics or AI Alignment?
Limited use · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Record details
Catalog link
Submitted location: CEUR-WS `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
- Ajay Vishwanath person
- Machine Ethics concept
- Machine Ethics or AI Alignment? paper
- Marija Slavkovik person
-
src-032 · Cited in 3 pages
Making Robots Conscious of Their Mental States
A logical-AI proposal describes machines reasoning about their own knowledge, ignorance and motivations.
Read Parallax writeup: Making Robots Conscious of Their Mental States
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Read source · Alternative version 1
Record details
Original 1995 report checked
Submitted location: AAAI `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. Original report inspected; distinguish later author web revisions.
Pages citing this source
- John McCarthy person
- Machine Introspection concept
- Making Robots Conscious of Their Mental States paper
-
src-033 · Cited in 3 pages
Making Robots Conscious of Their Mental States
A logical-AI proposal describes machines reasoning about their own knowledge, ignorance and motivations.
Read Parallax writeup: Making Robots Conscious of Their Mental States
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Record details
Catalog link
Submitted location: AAAI `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Likely duplicate: src-032
Pages citing this source
- John McCarthy person
- Machine Introspection concept
- Making Robots Conscious of Their Mental States paper
-
src-034 · Cited in 5 pages
Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework
Organizes evidence for claims that internally deployed AI will not defeat oversight through human manipulation.
Retain with stated limits · Reviewed 2026-10-09
Review reason and evidence
Retain; scoped original sections reviewed in batch 48, with benchmark/intent and proposed evaluation-to-deployment assumptions explicit.
- Methods, objectives and explicit trust/truthfulness assumptions checked
- Matching institutional manuscript, synthetic results and absence of convergence proof; OpenReview blocked
- Original safety-case arguments, limits, contributor roles and submission date
- Selected original v1 sections 3.1–3.3; rendered PDF page 13 control-argument assumptions checked. Direct/indirect inability, proposed mitigations, proxy bounds and insufficient trustworthiness evidence inspected. Full threat taxonomy/appendices, cited studies and replication not audited.
Record details
Catalog link
Submitted location: arXiv `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
-
src-035 · Cited in 2 pages
MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models
A benchmark tests coordinate-based navigation in English and Icelandic.
Read Parallax writeup: MazeEval
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Relevant original benchmark; small fixed maze sets, corrected statistics and exploratory large-maze trials constrain inference
Record details
Catalog link
Submitted location: arXiv `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
- Hafsteinn Einarsson person
- MazeEval paper
-
src-036 · Cited in 3 pages
Minding the moving self: neurodynamics of the self and psychotherapeutic implications
A theoretical review connects bodily movement, self-modeling, and psychotherapy.
Read Parallax writeup: Minding the Moving Self
Limited use · Reviewed 2026-10-07
Review reason and evidence
Interpretive human embodiment review; cannot establish AI consciousness or clinical treatment effect
Record details
Catalog link
Submitted location: Frontiers `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
- Embodied Cognition concept
- Minding the Moving Self paper
- Sharon Vaisvaser person
-
src-037 · Cited in 6 pages
Natural emergent misalignment from reward hacking in production RL
Controlled coding experiments link learned reward hacks to broader harmful behavior, with context-dependent limits on safety-training mitigations.
Read Parallax writeup: Natural Emergent Misalignment from Reward Hacking in Production RL
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
- Original v1 HTML methods, ablations and limitations; release and submission dates distinguished
- Batch 57: original arXiv v1 introduction/limitations, methods, section 3.1.2 sabotage and 3.1.3 mixture, sections 4.1–4.3/Figures 24–33 captions and discussion inspected. Targeted-training validation overlap, timing of inoculation, residual agentic failures, filtering/distillation and possible distribution mechanism checked. HTML figure captions assessed; numerical figure plots, complete appendices, code/data, revision comparison and independent replication not audited. Existing NotebookLM original ready; indexed identity/core prose inspected, completeness/fidelity not certified.
Record details
Catalog link
Submitted location: arXiv `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
- Emergent Misalignment concept
- Inoculation Prompting concept
- Key Papers: A Reading Guide note
- Monte MacDiarmid person
- Natural Emergent Misalignment from Reward Hacking in Production RL paper
- Reward Hacking concept
-
src-038 · Uncited
Producing Digital Reflections of Reality through Intercultural Filmmaking
Limited use · Reviewed 2026-10-07
Review reason and evidence
Practice-based filmmaking dissertation; institutional abstract supports identity and topic, not AI alignment claims
Record details
Catalog link
Submitted location: LSU Scholarly Repository `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-039 · Cited in 2 pages
Psychiatric Neural Networks and Precision Therapeutics by Machine Learning
A review examines clinical prediction and the challenges of translating machine learning into psychiatry.
Read Parallax writeup: Psychiatric Neural Networks and Precision Therapeutics by Machine Learning
Limited use · Reviewed 2026-10-07
Review reason and evidence
Clinical machine-learning review, not one independently validated diagnostic intervention
Record details
Catalog link
Submitted location: MDPI `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
-
src-041 · Cited in 5 pages
STA-RLHF: Stackelberg Aligned Reinforcement Learning with Human Feedback
Models coupled language-model and preference-model training as a leader–follower game.
Read Parallax writeup: STA-RLHF: Stackelberg Aligned Reinforcement Learning with Human Feedback
Limited use · Reviewed 2026-10-09
Review reason and evidence
Synthetic leader-follower RLHF experiments with oracle-labeled proxy objectives; manuscript versions differ and reward/divergence improvements are not uniform across them. Finite preference assumptions, leader-favorable ties and no convergence/deployment guarantee.
- Brown manuscript title/authors, Introduction p. 2 quotation, Experiments and Conclusion compared with separately retrieved RLC-headed original text; versions differ.
- Existing imported original PDF text: header RLJ / RLC 2024, five authors, methods, experiments and conclusion inspected. Exact PDF fidelity and revision identity unverified; public endpoint still challenged.
- Batch 51: Brown original selected sections 2–7 and Algorithm 1, rendered PDF page 7/Figure 1 inspected. Finite/complete/transitive preferences, leader-favorable ties, oracle labeling and nested versus simultaneous updates; classifier/word-count proxies and reward/KL comparison. Full proofs/appendices, code and replication not audited. Earlier RLC-headed extraction scope, differing tasks and unresolved revision identity preserved; no new RLC PDF verification claimed.
Read source · Alternative version 1
Record details
Verified identity
Submitted location: OpenReview `[PDF]`
Brown manuscript and separately retrieved RLC 2024-headed original text differ in terminology and experiments; title/five authors match. Exact revision identity, PDF fidelity and publication date remain unverified.
Pages citing this source
- Jacob Makar-Limanov person
- Outer Alignment concept
- Preference Aggregation concept
- STA-RLHF: Stackelberg Aligned Reinforcement Learning with Human Feedback paper
- Stackelberg Game concept
-
src-042 · Cited in 3 pages
Some Expert Systems Need Common Sense
McCarthy examines how narrow expertise can omit consequences, changing circumstances and knowledge of its own limits.
Read Parallax writeup: Some Expert Systems Need Common Sense
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Historical argument; author transcription reviewed, journal facsimile/cited evaluations unverified.
- Full author-hosted transcription including knowledge, reasoning and questions
- Publisher publication month checked
- Introduction: ontology, scope and illustrative cholera example; not observed clinical incident.
- Knowledge/reasoning distinction and richer ontology.
- Recommendations versus prognosis, concurrent events, knowledge and agents.
- Nonmonotonic reasoning, bird-cage communication example and discussion uncertainty about MYCIN needing common sense.
- Publisher November 1984, volume 426, pages 129–137; earliest release unverified.
Read source · Alternative version 1
Record details
Author transcription sections and publisher publication metadata checked; journal facsimile not audited
Submitted location: Cogprints / University of Southampton `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. Author-hosted transcription checked; 1984 publication distinguished from web conversion timestamp. Batch 64: author sections/discussion reread; November 1984 chronology verified. NotebookLM indexed entry contains introduction only, not complete original.
Pages citing this source
- Common-Sense Reasoning concept
- Distributional Shift concept
- Some Expert Systems Need Common Sense paper
-
src-043 · Uncited
THE DIZZINESS OF RECOGNITION: EXILE AS AN EDUCATIVE ENGAGEMENT by Parmis Aslanimehr
Limited use · Reviewed 2026-10-08
Review reason and evidence
Theoretical dissertation on exile, recognition and classroom education; abstract and preface support background use only. It supplies no empirical research design and does not establish AI alignment claims. Full quality review deferred as peripheral to this wiki.
Record details
Catalog link
Submitted location: Open Collections `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-044 · Cited in 2 pages
The Automated but Risky Game: Modeling Agent-to-Agent Negotiations and Transactions in Consumer Markets
Simulated buyer–seller transactions expose bargaining disparities and constraint violations.
Read Parallax writeup: The Automated but Risky Game
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Relevant original simulation with explicit model-estimated costs, judge and prompt limits; does not establish real consumer losses
Read source · Alternative version 1
Record details
Original full text checked
Submitted location: ResearchGate `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. Original v4 inspected; model-estimated wholesale costs and simulation limits noted.
Pages citing this source
- Shenzhe Zhu person
- The Automated but Risky Game paper
-
src-045 · Cited in 4 pages
The Challenge of Value Alignment: from Fairer Algorithms to AI Safety
Connects technical AI safety with fairness, participatory design and the plurality of social values.
Read Parallax writeup: The Challenge of Value Alignment: from Fairer Algorithms to AI Safety
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Read source · Alternative version 1
Record details
Verified identity
Submitted location: arXiv `[PDF]`
Original authors/title and identity checked October 7, 2026. Claim-review scope is recorded in the paper writeup and private research ledger.
Pages citing this source
-
src-046 · Cited in 2 pages
The Golem and the Game of Automation
A historical interpretation connects Wiener’s warnings with interactive learning.
Read Parallax writeup: The Golem and the Game of Automation
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Inspectable author-uploaded historical analysis; retain as interpretation rather than modern technical proof
Record details
Catalog link
Submitted location: ResearchGate `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
- Avery Slater person
- The Golem and the Game of Automation paper
-
src-047 · Cited in 2 pages
The Illusion of Rationality: Tacit Bias and Strategic Dominance in Frontier LLM Negotiation Games
Bargaining simulations test whether stronger language models negotiate consistently and fairly.
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Inspectable original simulation study relevant to agentic inequality; v1-specific conclusions and external-validity limits required
Record details
Catalog link
Submitted location: arXiv `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
-
src-048 · Uncited
The Neurobiology of Moral Behavior: Review and Neuropsychiatric Implications
Limited use · Reviewed 2026-10-07
Review reason and evidence
Human neurobiology review can provide historical background, not normative truth or evidence that AI has moral cognition
Record details
Catalog link
Submitted location: ResearchGate `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-049 · Cited in 1 page
The mirror mechanism and its potential role in autism spectrum disorder
Limited use · Reviewed 2026-10-07
Review reason and evidence
Human motor cognition review and debated autism hypothesis; avoid universal deficit claims and transfer to AI without evidence
Record details
Catalog link
Submitted location: ResearchGate `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
- Embodied Cognition concept
-
src-050 · Cited in 5 pages
TRAINING SOCIALLY ALIGNED LANGUAGE MODELS IN SIMULATED HUMAN SOCIETY
Uses simulated peer ratings, feedback and response revision as material for language-model alignment.
Read Parallax writeup: Training Socially Aligned Language Models in Simulated Human Society
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Record details
Catalog link
Submitted location: ICLR Proceedings `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
-
src-051 · Cited in 4 pages
Using the Veil of Ignorance to align AI systems with principles of justice
Tests how withholding knowledge of personal advantage affects choices of principles for an AI assistant.
Read Parallax writeup: Using the Veil of Ignorance to align AI systems with principles of justice
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Read source · Alternative version 1
Record details
Verified identity
Submitted location: ResearchGate `[URL]`
Original authors/title and identity checked October 7, 2026. Claim-review scope is recorded in the paper writeup and private research ledger.
Pages citing this source
-
src-052 · Cited in 2 pages
Veil-of-ignorance reasoning favors the greater good
Experiments test how impartial-perspective reasoning changes judgments in moral dilemmas.
Read Parallax writeup: Veil-of-Ignorance Reasoning Favors the Greater Good
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original controlled moral-judgment experiments; useful for procedure effects, not ethical truth
Read source · Alternative version 1
Record details
Original full text checked
Submitted location: ResearchGate `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. Original PNAS paper inspected; judgments do not establish normative correctness.
Pages citing this source
-
src-053 · Cited in 3 pages
Why Do Some Language Models Fake Alignment While Others Don't?
Compares training–deployment compliance gaps across 25 models and investigates why those gaps differ.
Read Parallax writeup: Why Do Some Language Models Fake Alignment While Others Don't?
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Abhay Sheshadri et al.
Read source · Alternative version 1
Record details
Verified identity
Submitted location: arXiv `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. Checked 2026-10-07.
Pages citing this source
-
src-054 · Cited in 4 pages
Dynamic value alignment through preference aggregation of multiple objectives
Combines multiple reinforcement-learning objectives with a changing voting population in a simulated traffic junction.
Read Parallax writeup: Dynamic value alignment through preference aggregation of multiple objectives
Retain with stated limits · Reviewed 2026-10-09
Review reason and evidence
Inspectible original traffic-simulation proof of concept; fixed individual preferences and changing voting population, two designer-supplied objectives and setting-dependent compromises. Truthfulness/trusted-rule assumptions; no demonstrated public legitimacy or real traffic safety.
- Methods, objectives and explicit trust/truthfulness assumptions checked
- Matching institutional manuscript, synthetic results and absence of convergence proof; OpenReview blocked
- Original safety-case arguments, limits, contributor roles and submission date
- Batch 51: original v1 selected sections 1, 3–6; rendered PDF page 6 Equations 1–7 inspected. Fixed individual preferences/changing voting population, upstream-only votes, normalized Q-score integration, 100 runs/six demands, half-half split and setting-dependent results; centralized/truthful/trusted-rule and designer-supplied objective assumptions. Code, replication, cited studies and complete revision comparison not audited.
Marcin Korecki, Damian Dailisan, Cesare Carissimo
Read source · Alternative version 1
Record details
Verified identity
Submitted location: arXiv `[PDF]`
Title and authors verified against the arXiv abstract record; previously imported as an arXiv identifier. v1 distinguishes fixed individual preferences from changing voter composition. Earlier review entries for Brown/Manipulation are preserved cross-source provenance, not evidence for traffic claims. Checked 2026-10-07.
Pages citing this source
-
src-055 · Cited in 6 pages
Aligned with Whom? Direct and Social Goals for AI Systems
Operator success and social welfare require different alignment and governance questions.
Read Parallax writeup: Aligned with Whom? Direct and Social Goals for AI Systems
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Record details
Original manuscript checked
Submitted location: Stanford Digital Economy `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. Same title and authors verified in the Stanford manuscript; duplicate of src-016.
Likely duplicate: src-016
Pages citing this source
- Aligned with Whom? Direct and Social Goals for AI Systems paper
- Anton Korinek person
- Avital Balwit person
- Direct Alignment concept
- Key Papers: A Reading Guide note
- Social Alignment concept
-
src-056 · Cited in 2 pages
“They parted illusions—they parted disclaim marinade”: Misalignment as structural fidelity in LLMs
A philosophical essay offers a linguistic interpretation of published safety evaluations.
Read Parallax writeup: Misalignment as Structural Fidelity in LLMs
Limited use · Reviewed 2026-10-07
Review reason and evidence
Philosophical interpretation, not controlled causal evidence; author acknowledges unfalsifiability risk
Mariana Lins Costa
Read source · Alternative version 1
Record details
Verified identity
Submitted location: arXiv `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. Checked 2026-10-07.
Pages citing this source
-
src-137 · Cited in 2 pages
Artificial intelligence in positive mental health: a narrative review
A narrative review discusses applications, cultural context, and validation needs.
Read Parallax writeup: Artificial Intelligence in Positive Mental Health
Limited use · Reviewed 2026-10-07
Review reason and evidence
Narrative review without pooled efficacy estimate; broad benefits require original intervention evidence and population scope
Record details
Catalog link
Submitted location: PMC / PubMed `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
-
src-138 · Cited in 2 pages
Beginnings of Artificial Intelligence in Medicine (AIM)
A participant’s historical commentary traces biomedical AI and its unresolved challenges.
Read Parallax writeup: Beginnings of Artificial Intelligence in Medicine
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original historical participant commentary; retain scoped to historical perspective rather than contemporary medical efficacy
Record details
Catalog link
Submitted location: PMC / PubMed `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
-
src-139 · Cited in 2 pages
De-humanizing Care: An Ethnography of Mental Health AI
An ethnography examines chatbot care, labor, and institutional incentives.
Read Parallax writeup: De-humanizing Care
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original situated ethnography relevant to care labor and incentives; interpretations and user accounts do not establish industry prevalence or clinical efficacy
Record details
Catalog link
Submitted location: UC Berkeley `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
- De-humanizing Care paper
- Valerie E. Black person
-
src-150 · Cited in 2 pages
Longtermist Political Philosophy: An Agenda for Future Research
A research agenda examines distant-future concerns alongside political values.
Read Parallax writeup: Longtermist Political Philosophy
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Relevant inspectable normative agenda; ethical arguments and speculative population scenarios must not be represented as measured facts
Record details
Catalog link
Submitted location: Oxford Academic `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
- Andreas T. Schmidt person
- Longtermist Political Philosophy paper
-
src-154 · Cited in 2 pages
Surfing the AI waves: historical evolution of AI in management
A historical framework organizes AI adoption and management research into five waves.
Read Parallax writeup: Surfing the AI Waves
Limited use · Reviewed 2026-10-07
Review reason and evidence
Interpretive historical wave framework; acknowledged selection/oversimplification limits and coined-in-1956 claim inconsistent with original 1955 proposal
Record details
Catalog link
Submitted location: Emerald Insight `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
- Matteo Cristofaro person
- Surfing the AI Waves paper
-
src-159 · Cited in 2 pages
Who's in the mirror: shaping organizational identity through AI
An abstract-level account of a conceptual paper on AI and organizational identity.
Read Parallax writeup: Who’s in the Mirror
Limited use · Reviewed 2026-10-07
Review reason and evidence
Publisher abstract only; conceptual identity claims cannot support causal effect estimates; full analysis inaccessible
Record details
Catalog link
Submitted location: Emerald Publishing `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
- Aslıhan Canbul Yaroğlu person
- Who’s in the Mirror paper
-
src-160 · Uncited
Working Paper | AI and the New Colonialism of Climate Data in the Global South
Review for removal · Reviewed 2026-10-07
Review reason and evidence
Broad claims lack original causal evaluation, and the statement that AI powers ND-GAIN is not supported by its documented indicator-averaging methodology. Recommend replacing factual use with original system documentation; retain ID pending editorial decision.
Record details
Catalog link
Submitted location: Public Affairs Research Institute (PARI) `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
notebook-662e96ae-43bd-4b12-85e5-0650f5ce7f4e · Uncited
[2508.16705] Assessing Consciousness-Related Behaviors in Large Language Models Using the Maze Test - arXiv
Unreviewed
Record details
Catalog link
Additional source discovered in the NotebookLM catalog, absent from the submitted list. Link recovered from source metadata.
-
ref-debate-2018 · Cited in 7 pages
AI safety via debate
Proposes competing AI arguments as a way for people to evaluate answers they could not produce or check unaided.
Read Parallax writeup: AI safety via debate
Retain with stated limits · Reviewed 2026-10-09
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
- Original paper HTML and arXiv identity/version history; conceptual scope and limitations checked
- Selected arXiv v2 original sections 2.2, 3.1–3.2/Table 2 and limitations; identity and first submission May 2/revision October 22, 2018 checked. Fixed classifier, true/false roles, nonzero pixel evidence, MCTS access and precommitment distinction inspected. Full proofs, complete v1/v2 comparison, code, cited studies and replication not audited. No matching 2018 original in current 180-source NotebookLM catalog; existing registered source, no new wiki source/import in this batch.
Record details
Verified identity
Primary arXiv record checked 2026-10-07; first submission May 2, 2018. Added following references in the Parallax NotebookLM corpus.
Pages citing this source
- AI safety via debate paper
- AI Safety via Debate concept
- Dario Amodei person
- Key Papers: A Reading Guide note
- Paul Christiano person
- Scalable Oversight concept
- Who Supplies Alignment Feedback? note
-
ref-cd3035dbef6c · Cited in 11 pages
Concrete Problems in AI Safety
Turns unintended machine-learning behavior into five practical research problems, from reward hacking to safe exploration.
Read Parallax writeup: Concrete Problems in AI Safety
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain: selected v2 baseline, reward-hacking and sparse-feedback proposals checked; earlier review scope preserved. Proposed experiments are not results.
- Original paper HTML and arXiv identity/version history; conceptual scope and limitations checked
- Selected original v2 PDF sections 2, 6 and 7; title/version checked. No independent audit of cited studies or proposed-experiment replication.
- Selected original arXiv v2 sections 3–5 read: impact baselines/empowerment caveats, reward-hacking mechanisms and proposed defenses, semi-supervised RL and proposed experiments. PDF pages 5 and 13 rendered and checked. Earlier sections 2/6/7 scopes preserved. No complete cited-study/proof/code audit, v1/v2 comparison or independent replication. Complete 186-source NotebookLM catalog audited before existing original PDF import; no matching title/identifier. Uploaded original ready/renamed cee57d3a-733b-40b1-a62b-fc04eed6eaf7; indexed six-author/title identity checked, full extraction fidelity not certified.
Record details
Verified identity
Title and identity checked against the original arXiv record. Checked 2026-10-07.
Pages citing this source
- 2023 pause letter on giant AI experiments event
- 2023 statement on AI risk event
- AI Alignment concept
- Concrete Problems in AI Safety paper
- Dario Amodei person
- Distributional Shift concept
- Key Papers: A Reading Guide note
- Outer Alignment concept
- Paul Christiano person
- Reward Hacking concept
- Scalable Oversight concept
-
ref-c4858d4ef280 · Cited in 8 pages
Risks from Learned Optimization in Advanced Machine Learning Systems
Separates the objective used to train a model from the objective a learned optimizer might pursue.
Read Parallax writeup: Risks from Learned Optimization in Advanced Machine Learning Systems
Retain with stated limits · Reviewed 2026-10-09
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Record details
Verified identity
Title and identity checked against the original arXiv record. Checked 2026-10-07.
Pages citing this source
- Deceptive Alignment concept
- Distributional Shift concept
- Evan Hubinger person
- Inner Alignment concept
- Key Papers: A Reading Guide note
- Mesa-Optimization concept
- Outer Alignment concept
- Risks from Learned Optimization in Advanced Machine Learning Systems paper
-
ref-030a9cc7079a · Cited in 4 pages
Artificial Intelligence, Values and Alignment
Distinguishes alignment targets and argues for fair principles that can receive endorsement despite moral disagreement.
Read Parallax writeup: Artificial Intelligence, Values and Alignment
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Record details
Verified identity
Title and identity checked against the original arXiv record. Checked 2026-10-07.
Pages citing this source
- AI Alignment concept
- Artificial Intelligence, Values and Alignment paper
- Iason Gabriel person
- Key Papers: A Reading Guide note
-
ref-ad43ed014e22 · Cited in 5 pages
Society-in-the-Loop: Programming the Algorithmic Social Contract
Connects human oversight with stakeholder negotiation and monitoring of an algorithmic social contract.
Read Parallax writeup: Society-in-the-Loop: Programming the Algorithmic Social Contract
Retain with stated limits · Reviewed 2026-10-07
Review reason and evidence
Original material checked during the initial academic pass; cite within the methods, argument and limits stated in the linked writeup.
Record details
Verified identity
Title and identity checked against the original arXiv record. Checked 2026-10-07.
Pages citing this source
-
ref-e178cbf073c0 · Cited in 6 pages
Categorizing Variants of Goodhart's Law
Four mechanisms explain why greater optimization of a proxy can undermine its intended goal.
Read Parallax writeup: Categorizing Variants of Goodhart's Law
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain conceptual taxonomy with explicit selection/intervention/overlap distinctions;not empirical frequency or universal mitigation evidence.
- Original v4 HTML taxonomy and arXiv submission/revision history
- Batch72: original arXiv v4 PDF pp1–10 main text/equations/examples/conclusion read; references inspected as citations only. Original identity/authors/v1 submission2018-03-13/v4 revision2019-02-24 checked;PDF printed2019-02-26 and experimental HTML displayed2026-08-24 not used as release dates. Rendered p2 exact12-word overlap quotation checked. Cited originals,earlier Garrabrant post,version differences,proof/empirical validation and full extraction fidelity unreviewed. Indexed original title/authors/selected taxonomy prose checked.
Record details
Verified identity
Title and identity checked against the original arXiv record. Checked 2026-10-07.
Pages citing this source
- Categorizing Variants of Goodhart's Law paper
- David Manheim person
- From objectives to accountable control note
- Goodhart's Law concept
- Key Papers: A Reading Guide note
- Scott Garrabrant person
-
ref-28f9d1d93970 · Cited in 4 pages
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
Demonstrates designed selective underperformance and capability concealment in language-model evaluations.
Read Parallax writeup: AI Sandbagging: Language Models can Strategically Underperform on Evaluations
Retain with stated limits · Reviewed 2026-10-09
Review reason and evidence
Retain; scoped original sections reviewed in batch 48, with benchmark/intent and proposed evaluation-to-deployment assumptions explicit.
- Comparison, interventions, hypothesis and fine-tuning limitations, submission and author roles
- Designed prompting/password experiments; source-version scope recorded
- Constructed prompt/document/training methodology and inference limits; release date checked
- Selected original v3 sections 2–5 and 7–8; rendered PDF page 6 Figures 4–5/Table 1 checked; definition versus tested capability, refusal exclusions, password generalization and emulation distinguished. Full appendices/code, revision comparison and replication not audited.
Record details
Verified identity
Earlier reference retained to preserve its source anchor; linked to the canonical paper record.
Likely duplicate: src-011
Pages citing this source
-
src-169 · Cited in 2 pages
Circumscription—A Form of Nonmonotonic Reasoning
McCarthy formalizes revisable assumptions for planning when action conditions cannot all be enumerated.
Read Parallax writeup: Circumscription—A Form of Nonmonotonic Reasoning
Retain with stated limits · Reviewed 2026-10-09
Review reason and evidence
Retain as a historical formal proposal for defeasible planning; selected author reprint sections reviewed, not validated modern safety method. Distinguish 1980 publication from later reprint/addendum and superseding 1986 paper.
- Author-hosted reprint: abstract and sections 1–4 selected passages; not complete mathematical review. Header 1986; later addendum not used as 1980 evidence.
- Puzzle versus real-world conjecture and revision when a bridge is added.
- Remarks 1, 4 and 7: first-order logic, heuristic retraction and representation dependence.
- Author bibliographic record: Artificial Intelligence 13 (1980), 27–39; later addendum and superseding 1986 formalism warning.
Record details
Author-hosted reprint selected sections checked
John McCarthy; Artificial Intelligence 13 (1980), 27–39. Author-hosted PDF header says 1986 and includes a later addendum; publisher original facsimile not accessed. Author recommends the superseding 1986 Applications paper for the later formalism.
Pages citing this source
-
src-170 · Cited in 3 pages
Analyzing and Improving Chain-of-Thought Monitorability Through Information Theory
An information-theoretic analysis and controlled training experiments test ways to preserve useful signals in reasoning traces.
Retain with stated limits · Reviewed 2026-10-09
Review reason and evidence
Retain as a formal analysis and controlled training study; monitorability proxy is not sufficient in general and interventions do not eliminate all hacking.
- Original PDF sections 3–6, Figure 3 caption and selected Appendix B.1–B.3 text; formal statements and experimental setup/results/limitations inspected, not full proof review or replication.
- Title, authors, version and first arXiv submission date checked; no inferred first-public-availability or peer-review claim.
Usman Anwar; Tim Bakker; Dana Kianfar; Cristina Pinneri; Christos Louizos
Read source · Alternative version 1
Record details
Original PDF selected sections checked
2026 arXiv v1 preprint. First submission February 20; PDF also prints February 23, 2026. Original PDF sections 3–6, Figure 3 caption and selected Appendix B.1–B.3 text; formal statements and experimental setup/results/limitations inspected, not full proof review or replication.
Pages citing this source
-
src-171 · Cited in 2 pages
Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings
Matched prompt interventions show that contextual nudges can change answers while escaping a reasoning-trace monitor.
Read Parallax writeup: Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings
Retain with stated limits · Reviewed 2026-10-09
Review reason and evidence
Retain as controlled evidence of context-sensitive monitoring failures; LLM-judge detection on behavior-shift cases, not real-world incident prevalence or undetectability in principle.
- Original PDF sections 1, 3–6 and Figure 1 summary; matched intervention design, conditional detection, API trace access and LLM-judge limitations inspected; full appendices and code not audited.
- Title, authors, version and first arXiv submission date checked; no inferred first-public-availability or peer-review claim.
Agatha Duzan; Asa Cooper Stickland
Read source · Alternative version 1
Record details
Original PDF selected sections checked
2026 arXiv v1 preprint. First submission August 5, 2026. Original PDF sections 1, 3–6 and Figure 1 summary; matched intervention design, conditional detection, API trace access and LLM-judge limitations inspected; full appendices and code not audited.
Pages citing this source
-
src-172 · Cited in 3 pages
Algorithms for Inverse Reinforcement Learning
Ng and Russell infer reward functions that make observed decisions optimal, while exposing ambiguity in that inference.
Read Parallax writeup: Algorithms for Inverse Reinforcement Learning
Retain with stated limits · Reviewed 2026-10-09
Review reason and evidence
Retain as historical objective-inference research; compatible rewards are underdetermined, selection heuristics are not verified recovery of human values.
- Original author-hosted PDF rendered pages 1–4 and 7–8: formulation, degeneracy, selection heuristics and conclusion checked visually. Text extraction and NotebookLM indexed text garbled; ready original file copy verified, usable indexed identity/passages not verified. Full derivations/proofs and replication not audited.
- Author publication archive confirms title, authors, ICML venue and publication year; no day-specific date inferred.
Andrew Y. Ng; Stuart Russell
Record details
Original PDF selected sections checked
Original author-hosted PDF rendered pages 1–4 and 7–8: formulation, degeneracy, selection heuristics and conclusion checked visually. Text extraction and NotebookLM indexed text garbled; ready original file copy verified, usable indexed identity/passages not verified. Full derivations/proofs and replication not audited.
Pages citing this source
- AI Alignment concept
- Algorithms for Inverse Reinforcement Learning paper
- Stuart Russell person
-
src-173 · Cited in 3 pages
Apprenticeship Learning via Inverse Reinforcement Learning
Abbeel and Ng match demonstrated feature expectations to obtain comparable task performance without identifying the true reward.
Read Parallax writeup: Apprenticeship Learning via Inverse Reinforcement Learning
Retain with stated limits · Reviewed 2026-10-09
Review reason and evidence
Retain as historical task learning with conditional performance guarantees; no true-reward recovery or deployed safety claim.
- Original PDF sections 1–6 and Table 1 inspected: feature matching, policy selection/mixture, theorem assumptions, simulation and unsafe imitation limits. Ready NotebookLM original copy and indexed title/authors/selected passages verified. Appendix proofs, supplement, code and replication not audited.
- Author publication archive confirms title, authors, ICML venue and publication year; no day-specific date inferred.
Pieter Abbeel; Andrew Y. Ng
Record details
Original PDF selected sections checked
Original PDF sections 1–6 and Table 1 inspected: feature matching, policy selection/mixture, theorem assumptions, simulation and unsafe imitation limits. Ready NotebookLM original copy and indexed title/authors/selected passages verified. Appendix proofs, supplement, code and replication not audited.
Pages citing this source
- AI Alignment concept
- Apprenticeship Learning via Inverse Reinforcement Learning paper
- Outer Alignment concept
-
src-174 · Cited in 3 pages
Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback
A controlled repeated-task study finds monitoring awareness and rare stated concealment intent, without successful evasion.
Retain with stated limits · Reviewed 2026-10-09
Review reason and evidence
Retain as a controlled monitoring-awareness study; no successful evasion or independently invented harmful objectives demonstrated.
- Original arXiv v1 PDF main sections 1–6 and rendered Figure 1 inspected; assigned side tasks, episodic feedback, explicit-statement rubric, three failed concealment intents and zero evasion checked. Appendix rubric/transcript details, code and replication not audited. Ready NotebookLM original PDF copy and indexed identity/selected passages verified; extraction completeness not certified.
- Original arXiv title/authors/submission and revision identity checked; no peer-review claim.
Thomas Jiralerspong; Flemming Kondrup; Yoshua Bengio
Read source · Alternative version 1
Record details
Original PDF selected sections checked
Original arXiv v1 PDF main sections 1–6 and rendered Figure 1 inspected; assigned side tasks, episodic feedback, explicit-statement rubric, three failed concealment intents and zero evasion checked. Appendix rubric/transcript details, code and replication not audited. Ready NotebookLM original PDF copy and indexed identity/selected passages verified; extraction completeness not certified.
Pages citing this source
-
src-175 · Cited in 3 pages
Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute
A coding benchmark study finds complementary prompted and fine-tuned monitors improve detection, with adaptive-attacker robustness untested.
Read Parallax writeup: Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute
Retain with stated limits · Reviewed 2026-10-09
Review reason and evidence
Retain as scoped code-detection evidence; selected ensemble gains are not adaptive-attacker robustness or deployed safety guarantees.
- Original arXiv v2 PDF main sections 1–7, Tables 1–2 and selection/limitations inspected; test-set ranking distinguished from validation operating points and cross-dataset transfer. Full appendices, code, data provenance and replication not audited. Ready NotebookLM original PDF copy and indexed identity/selected passages verified; extraction completeness not certified.
- Original arXiv title/authors/submission and revision identity checked; no peer-review claim.
Eugene Koran; Yejun Yun; Samantha Tetef; Benjamin Arnav; Pablo Bernabeu-Pérez
Read source · Alternative version 1
Record details
Original PDF selected sections checked
Original arXiv v2 PDF main sections 1–7, Tables 1–2 and selection/limitations inspected; test-set ranking distinguished from validation operating points and cross-dataset transfer. Full appendices, code, data provenance and replication not audited. Ready NotebookLM original PDF copy and indexed identity/selected passages verified; extraction completeness not certified.
Pages citing this source
-
src-176 · Cited in 2 pages
Reinforcement Learning: A Survey
A 1996 survey distinguishes the objective an agent optimizes from the performance and penalties of learning.
Read Parallax writeup: Reinforcement Learning: A Survey
Retain with stated limits · Reviewed 2026-10-09
Review reason and evidence
Retain as historical general learning/control survey; objective criterion and learning-period penalties are distinct. Modern alignment relationship editorial.
- Original CMU-hosted JAIR PDF selected section 1.1–1.3 passages; rendered PDF pages 5–6 (printed 241–242), Figure 2 and learning-performance discussion inspected. Figure caption h=4 versus adjacent text h=5 discrepancy observed; no numeric example claim used. Full survey algorithms, cited experiments and replication not audited. Notebook original ready; identity and selected prose checked, glyph errors and completeness limitations retained.
- JAIR original title/authors/publication date May 1, 1996 verified; arXiv v1 same-day submission verified separately.
Leslie Pack Kaelbling; Michael L. Littman; Andrew W. Moore
Read source · Alternative version 1 · Alternative version 2
Record details
Original PDF selected sections checked
Original CMU-hosted JAIR PDF selected section 1.1–1.3 passages; rendered PDF pages 5–6 (printed 241–242), Figure 2 and learning-performance discussion inspected. Figure caption h=4 versus adjacent text h=5 discrepancy observed; no numeric example claim used. Full survey algorithms, cited experiments and replication not audited. Notebook original ready; identity and selected prose checked, glyph errors and completeness limitations retained.
Pages citing this source
- Outer Alignment concept
- Reinforcement Learning: A Survey paper
-
src-177 · Cited in 4 pages
Policy invariance under reward transformations: Theory and application to reward shaping
A 1999 reward-shaping result preserves optimal policies under stated conditions, while leaving the original objective to be justified.
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain as conditional formal reward-shaping result; neither guaranteed convergence speed nor validation of an aligned original objective.
- Author-linked Berkeley original PDF selected sections 1–3 and 5; rendered page 4 Theorem 1/Equation 2 checked independently. Stanford copy garbled; Berkeley prose usable but mathematical glyphs missing in extraction and notebook index. Full necessity/sufficiency proof audit, section 4 experiments and replication not performed. Bicycle anecdote attributed to this paper; underlying 1998 study not reviewed. Notebook original ready; indexed identity/selected prose verified, formula not established from index.
- Author bibliography records ICML-99, Bled, Slovenia, 1999; no exact first-release date inferred.
- Original section 4 and section 5 context inspected through rendered Stanford author-copy PDF pages 1 and 6–8, cross-checked against Berkeley web extraction. Figures 1–2, 40-run averaging, Sarsa settings, distance heuristic and ordered-subgoal state augmentation checked. Berkeley direct download returned 404 and web screenshots failed; registered Stanford alternate retrieved and visually legible, text extraction garbled. Existing NotebookLM original 4d8d12e1-99af-4c51-95c3-bc70ed86ca7b ready; title/three-author identity and selected indexed prose verified, mathematical extraction unreliable. Full proof audit, code, uncertainty analysis, underlying bicycle study, version equivalence and independent replication not performed. Earlier scopes preserved.
Andrew Y. Ng; Daishi Harada; Stuart Russell
Read source · Alternative version 1
Record details
Original selected sections and rendered experiment figures checked
Author-linked Berkeley original PDF selected sections 1–3 and 5; rendered page 4 Theorem 1/Equation 2 checked independently. Stanford copy garbled; Berkeley prose usable but mathematical glyphs missing in extraction and notebook index. Full necessity/sufficiency proof audit, section 4 experiments and replication not performed. Bicycle anecdote attributed to this paper; underlying 1998 study not reviewed. Notebook original ready; indexed identity/selected prose verified, formula not established from index. Additional experiment review: Original section 4 and section 5 context inspected through rendered Stanford author-copy PDF pages 1 and 6–8, cross-checked against Berkeley web extraction. Figures 1–2, 40-run averaging, Sarsa settings, distance heuristic and ordered-subgoal state augmentation checked. Berkeley direct download returned 404 and web screenshots failed; registered Stanford alternate retrieved and visually legible, text extraction garbled. Existing NotebookLM original 4d8d12e1-99af-4c51-95c3-bc70ed86ca7b ready; title/three-author identity and selected indexed prose verified, mathematical extraction unreliable. Full proof audit, code, uncertainty analysis, underlying bicycle study, version equivalence and independent replication not performed. Earlier scopes preserved.
Pages citing this source
-
src-178 · Cited in 7 pages
Safe Learning Under Irreversible Dynamics via Asking for Help
Turns unintended machine-learning behavior into five practical research problems, from reward hacking to safe exploration.
Read Parallax writeup: Concrete Problems in AI Safety · Read Parallax writeup: Safe Learning Under Irreversible Dynamics via Asking for Help
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain as conditional theoretical mentor-guided learning result; mentor-relative reward and query bounds do not establish absolute safety or practical deployment readiness.
- JMLR 2026 original PDF selected sections 1–3, 4.3, 5.2 and 6–7, Algorithm 3; rendered PDF pages 4 (Figure 2) and 17 (Theorems 9–10/Corollary 11) inspected. Mentor-relative separate-trajectory regret, policy-class/smoothness/local-generalization assumptions, binary illustrative construction, query diameter factor and practical limits checked. Full reduction proofs, appendices, cited studies and complete arXiv/JMLR version comparison not audited. Notebook original ready; indexed identity and selected prose checked, mathematical fidelity/completeness not certified.
- First submission February 19, 2025; September 11, 2026 v4 revision verified. JMLR original header records submitted 9/25, revised 4/26, published 6/26. Chronology uses first submission; no invented publication day.
- Additional October10 integration review: JMLR original identity/header, sections1–2.1, Algorithm3, selected section6.3 and section7 inspected; rendered p4 Figure2 and surrounding comparison text checked. Exact14-word section1.1/p2 quote verified against PDF extraction and ready NotebookLM indexed original. Mentor-relative objective, separate trajectories, action-conditioned familiarity, perfect-distance/full-observability limits rechecked. Full proofs, appendices, cited studies, arXiv/JMLR version equivalence and replication remain unreviewed; mathematical extraction/completeness not certified. Earlier scopes preserved.
Benjamin Plaut; Juan Liévano-Karim; Hanlin Zhu; Stuart Russell
Read source · Alternative version 1 · Alternative version 2
Record details
Original PDF selected sections checked
JMLR 2026 original PDF selected sections 1–3, 4.3, 5.2 and 6–7, Algorithm 3; rendered PDF pages 4 (Figure 2) and 17 (Theorems 9–10/Corollary 11) inspected. Mentor-relative separate-trajectory regret, policy-class/smoothness/local-generalization assumptions, binary illustrative construction, query diameter factor and practical limits checked. Full reduction proofs, appendices, cited studies and complete arXiv/JMLR version comparison not audited. Notebook original ready; indexed identity and selected prose checked, mathematical fidelity/completeness not certified. Additional October10 integration review: JMLR original identity/header, sections1–2.1, Algorithm3, selected section6.3 and section7 inspected; rendered p4 Figure2 and surrounding comparison text checked. Exact14-word section1.1/p2 quote verified against PDF extraction and ready NotebookLM indexed original. Mentor-relative objective, separate trajectories, action-conditioned familiarity, perfect-distance/full-observability limits rechecked. Full proofs, appendices, cited studies, arXiv/JMLR version equivalence and replication remain unreviewed; mathematical extraction/completeness not certified. Earlier scopes preserved.
Pages citing this source
- Concrete Problems in AI Safety paper
- Distributional Shift concept
- Outer Alignment concept
- Safe Learning Under Irreversible Dynamics via Asking for Help paper
- Safety Case concept
- Scalable Oversight concept
- Stuart Russell person
-
src-179 · Cited in 4 pages
Training language models to follow instructions with human feedback
A 2022 study uses demonstrations, human rankings and reinforcement learning to improve GPT-3 instruction following, while exposing limits of labeler preferences and safety.
Read Parallax writeup: Training language models to follow instructions with human feedback
Retain with stated limits · Reviewed 2026-10-09
Review reason and evidence
Retain as an original empirical instruction-following/RLHF study; preference gains are distribution- and evaluator-specific, not universal value alignment or safety guarantees.
- Batch 52: original arXiv v1 selected introduction, sections 3.1–3.6, 4.2 and 5.1–5.4; rendered PDF pages 3 (Figure 2/results) and 14 (Figure 7/toxicity and bias) inspected. Demonstration/ranking/PPO pipeline, KL penalty/pretraining mix, customer-disjoint mostly-English data, prompt-conditioned toxicity, selected-labeler target and harmful-instruction limitations. Full appendices, code, cited studies and replication not audited. Notebook original PDF ready; indexed title/authors and selected prose checked, full mathematical fidelity/completeness not certified.
- Title/authors and sole v1 submission March 4, 2022 verified; chronology uses submission, no invented release/publication date.
Long Ouyang; Jeff Wu; Xu Jiang; Diogo Almeida; Carroll L. Wainwright; Pamela Mishkin; Chong Zhang; Sandhini Agarwal; Katarina Slama; Alex Ray; John Schulman; Jacob Hilton; Fraser Kelton; Luke Miller; Maddie Simens; Amanda Askell; Peter Welinder; Paul Christiano; Jan Leike; Ryan Lowe
Read source · Alternative version 1
Record details
Original PDF selected sections checked
Batch 52: original arXiv v1 selected introduction, sections 3.1–3.6, 4.2 and 5.1–5.4; rendered PDF pages 3 (Figure 2/results) and 14 (Figure 7/toxicity and bias) inspected. Demonstration/ranking/PPO pipeline, KL penalty/pretraining mix, customer-disjoint mostly-English data, prompt-conditioned toxicity, selected-labeler target and harmful-instruction limitations. Full appendices, code, cited studies and replication not audited. Notebook original PDF ready; indexed title/authors and selected prose checked, full mathematical fidelity/completeness not certified.
Pages citing this source
-
src-180 · Cited in 2 pages
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
An automated-R&D benchmark tests artifact access and reasoning access for sabotage monitoring.
Read Parallax writeup: ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
Retain with stated limits · Reviewed 2026-10-09
Review reason and evidence
Retain as a controlled automated-R&D sabotage/monitoring evaluation; successful-attack conditional detection and limited strategic concealment do not establish deployment safety or incident prevalence.
- Selected arXiv v2 original sections 1 and 3–6, Tables 2–3; rendered PDF pages 5 (access/setup) and 10 (conditional detection/AUC tables and discussion) inspected. Successful-run conditioning, artifact/CoT access, data-carried sabotage, threshold/calibration and strategic-hiding/diffuse-threat limits checked. Full appendices, cited studies, v1/v2 comparison, code/data and replication not audited. Notebook original ready; indexed title/authors and selected setup/results/limitations verified, extraction completeness and mathematical fidelity not certified.
- Authors/title, first submission July 21, 2026 and reviewed v2 revision July 29, 2026 verified; no peer-reviewed publication claim.
Lena Libon; Ben Rank; Jehyeok Yeon; David Schmotz; Jeremy Qin; Daniel Donnelly; Derck Prinzhorn; Maksym Andriushchenko
Read source · Alternative version 1
Record details
Original PDF selected sections checked
Selected arXiv v2 original sections 1 and 3–6, Tables 2–3; rendered PDF pages 5 (access/setup) and 10 (conditional detection/AUC tables and discussion) inspected. Successful-run conditioning, artifact/CoT access, data-carried sabotage, threshold/calibration and strategic-hiding/diffuse-threat limits checked. Full appendices, cited studies, v1/v2 comparison, code/data and replication not audited. Notebook original ready; indexed title/authors and selected setup/results/limitations verified, extraction completeness and mathematical fidelity not certified.
Pages citing this source
-
src-181 · Cited in 3 pages
Constitutional AI: Harmlessness from AI Feedback
A 2022 experiment trains assistants to critique, revise and judge responses using written principles, improving human-rated harmlessness while retaining human helpfulness feedback.
Read Parallax writeup: Constitutional AI: Harmlessness from AI Feedback
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain as original controlled evidence for principle-guided self-revision and AI harmlessness feedback; human helpfulness feedback remains, evaluation preferences and proxy limits constrain generalization.
- Selected original arXiv v1 sections 1.1–1.4, 3.1–3.3, 4.1–4.5, 5–6.2 and Appendix C principles; rendered PDF pages 3 (Figure 2 tradeoff/evaluation instructions) and 14 (Figure 10/absolute-harmfulness setup) inspected. Hybrid human-helpfulness/AI-harmlessness labels, critique/revision and RL stages, different initialization/evaluation instructions, clamped probabilities, reward overoptimization and research-selected principles checked. Full appendices, underlying prior datasets, implementation and independent replication not audited. Notebook original PDF ready; indexed title/authors and opening prose checked; figures/index completeness not certified.
- Title, 51-author identity and v1 submission December 15, 2022 verified; no peer-reviewed publication assertion.
Yuntao Bai; Saurav Kadavath; Sandipan Kundu; Amanda Askell; Jackson Kernion; Andy Jones; Anna Chen; Anna Goldie; Azalia Mirhoseini; Cameron McKinnon; Carol Chen; Catherine Olsson; Christopher Olah; Danny Hernandez; Dawn Drain; Deep Ganguli; Dustin Li; Eli Tran-Johnson; Ethan Perez; Jamie Kerr; Jared Mueller; Jeffrey Ladish; Joshua Landau; Kamal Ndousse; Kamile Lukosuite; Liane Lovitt; Michael Sellitto; Nelson Elhage; Nicholas Schiefer; Noemi Mercado; Nova DasSarma; Robert Lasenby; Robin Larson; Sam Ringer; Scott Johnston; Shauna Kravec; Sheer El Showk; Stanislav Fort; Tamera Lanham; Timothy Telleen-Lawton; Tom Conerly; Tom Henighan; Tristan Hume; Samuel R. Bowman; Zac Hatfield-Dodds; Ben Mann; Dario Amodei; Nicholas Joseph; Sam McCandlish; Tom Brown; Jared Kaplan
Read source · Alternative version 1
Record details
Original PDF selected sections checked
Selected original arXiv v1 sections 1.1–1.4, 3.1–3.3, 4.1–4.5, 5–6.2 and Appendix C principles; rendered PDF pages 3 (Figure 2 tradeoff/evaluation instructions) and 14 (Figure 10/absolute-harmfulness setup) inspected. Hybrid human-helpfulness/AI-harmlessness labels, critique/revision and RL stages, different initialization/evaluation instructions, clamped probabilities, reward overoptimization and research-selected principles checked. Full appendices, underlying prior datasets, implementation and independent replication not audited. Notebook original PDF ready; indexed title/authors and opening prose checked; figures/index completeness not certified.
Pages citing this source
- Constitutional AI: Harmlessness from AI Feedback paper
- Outer Alignment concept
- Scalable Oversight concept
-
src-185 · Cited in 5 pages
Deep reinforcement learning from human preferences
Turns unintended machine-learning behavior into five practical research problems, from reward hacking to safe exploration.
Read Parallax writeup: Concrete Problems in AI Safety · Read Parallax writeup: Deep reinforcement learning from human preferences · Read Parallax writeup: Training language models to follow instructions with human feedback
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain as original empirical scaling of preference-based reward learning to deep RL; task/rater-specific simulated results and proxy failures do not establish universal values or deployment safety.
- Original arXiv v1 introduction/related-work, sections 2.1–2.2.4, 3.1–3.3 and 4 inspected; rendered PDF pages 5 (preference likelihood/feedback choices) and 8 (Figure 3 Atari curves/run counts and Figure 4/novel behaviors) checked. Human versus synthetic-oracle feedback, asynchronous reward/policy updates, single human-feedback runs, task-specific failures and offline-feedback Pong failure assessed. Section 3.1.2 synthetic-label-count prose conflicts with Figure 3; no exact synthetic-budget claim used. Full appendices, videos, implementation, cited originals, version comparison and independent replication not audited. Notebook original ready; indexed identity and ablation prose checked, complete extraction and mathematical fidelity not certified.
- Title/six authors and first submission June 12, 2017 verified; subsequent revisions July 2/13, 2017 and February 17, 2023 kept separate. HTML rendering date August 24, 2026 not used as research/publication date.
Paul F. Christiano; Jan Leike; Tom B. Brown; Miljan Martic; Shane Legg; Dario Amodei
Read source · Alternative version 1
Record details
Original PDF selected sections checked
Original arXiv v1 introduction/related-work, sections 2.1–2.2.4, 3.1–3.3 and 4 inspected; rendered PDF pages 5 (preference likelihood/feedback choices) and 8 (Figure 3 Atari curves/run counts and Figure 4/novel behaviors) checked. Human versus synthetic-oracle feedback, asynchronous reward/policy updates, single human-feedback runs, task-specific failures and offline-feedback Pong failure assessed. Section 3.1.2 synthetic-label-count prose conflicts with Figure 3; no exact synthetic-budget claim used. Full appendices, videos, implementation, cited originals, version comparison and independent replication not audited. Notebook original ready; indexed identity and ablation prose checked, complete extraction and mathematical fidelity not certified.
Pages citing this source
-
src-186 · Cited in 4 pages
How does information access affect LLM monitors' ability to detect sabotage?
A coding benchmark study finds complementary prompted and fine-tuned monitors improve detection, with adaptive-attacker robustness untested.
Read Parallax writeup: Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute · Read Parallax writeup: How does information access affect LLM monitors' ability to detect sabotage?
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain as original empirical evidence about information filtering in retrospective sabotage monitoring; conditional benchmark detection does not establish prospective intervention or deployment safety.
- Original arXiv v1 selected introduction, sections 3–4, 5.1–5.2, 5.4, 5.6 and 6 inspected; rendered PDF page 7 Figure 4 and averaging/baseline/evaluator passages checked. Conditional successful-attack filtering, asynchronous retrospective monitoring, separate contexts, variable awareness conditions and extractor/evaluator failure modes assessed. Full appendices, code/data provenance, model-version audit, v2 comparison and independent replication not reviewed. Notebook original ready and indexed identity checked; complete extraction/math fidelity not certified.
- Five authors/title, first submission January 28, 2026 and v2 revision February 5 verified; v1 selected review, no version comparison.
Rauno Arike; Raja Mehta Moreno; Rohan Subramani; Shubhorup Biswas; Francis Rhys Ward
Read source · Alternative version 1
Record details
Original PDF selected sections checked
Original arXiv v1 selected introduction, sections 3–4, 5.1–5.2, 5.4, 5.6 and 6 inspected; rendered PDF page 7 Figure 4 and averaging/baseline/evaluator passages checked. Conditional successful-attack filtering, asynchronous retrospective monitoring, separate contexts, variable awareness conditions and extractor/evaluator failure modes assessed. Full appendices, code/data provenance, model-version audit, v2 comparison and independent replication not reviewed. Notebook original ready and indexed identity checked; complete extraction/math fidelity not certified.
Pages citing this source
-
src-187 · Cited in 4 pages
Learning to summarize from human feedback
A 2020 summarization study learns rewards from human comparisons, improving judged quality while showing how stronger optimization can exploit the learned proxy.
Read Parallax writeup: Learning to summarize from human feedback · Read Parallax writeup: Training language models to follow instructions with human feedback
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain as original human-feedback summarization experiment; selected evaluator preferences and transfer findings do not establish universal values, reliable factuality or deployment safety.
- Original arXiv v1 introduction, sections 3–5 and broader-impacts discussion inspected; rendered PDF pages 8 (Figure 5 early reward-model overoptimization) and 9 (Figure 7 rejection sampling and unequal human-data baseline limitation) checked. Task/rater-specific judgments, summary-length confound, news transfer and learned-proxy limits assessed. PDF extraction reconstructed a broken xref; selected rendered pages verified, full extraction fidelity not certified. Full appendices, code/data, cited originals, revision comparison and independent replication not audited. Complete 164-file wiki/195-register and 187-source notebook duplicate audit found no original match. Imported original ready/renamed; indexed title/nine-author identity checked.
- Title/nine authors and first submission September 2, 2020 verified; October 27, 2020 v2 and February 15, 2022 v3 kept separate, not compared.
Nisan Stiennon; Long Ouyang; Jeff Wu; Daniel M. Ziegler; Ryan Lowe; Chelsea Voss; Alec Radford; Dario Amodei; Paul Christiano
Read source · Alternative version 1
Record details
Original PDF selected sections checked
Original arXiv v1 introduction, sections 3–5 and broader-impacts discussion inspected; rendered PDF pages 8 (Figure 5 early reward-model overoptimization) and 9 (Figure 7 rejection sampling and unequal human-data baseline limitation) checked. Task/rater-specific judgments, summary-length confound, news transfer and learned-proxy limits assessed. PDF extraction reconstructed a broken xref; selected rendered pages verified, full extraction fidelity not certified. Full appendices, code/data, cited originals, revision comparison and independent replication not audited. Complete 164-file wiki/195-register and 187-source notebook duplicate audit found no original match. Imported original ready/renamed; indexed title/nine-author identity checked.
Pages citing this source
-
src-188 · Cited in 4 pages
Diffuse AI Control on Fuzzy Tasks
A prompt-based adversarial experiment tests whether weak judges can reward poor research proposals and whether judging prompts can resist adaptive attacks.
Read Parallax writeup: Diffuse AI Control on Fuzzy Tasks · Read Parallax writeup: ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain as original 2026 adversarial oversight experiment; prompted score divergence and tested defenses do not establish spontaneous scheming, successful real-world sabotage or general training/deployment guarantees.
- Original arXiv v2 sections 1–5 and 7–8 inspected; rendered PDF pages 5 (Figure 4 and scorer setup) and 6 (attack slack and Figure 5 defense) checked. Prompt optimization versus hypothetical fine-tuning, model-based proxy versus ground truth, adaptive rubric access, defense discovery affordances and task/distribution limits assessed. Full appendices, code/data, cited originals, v1/v2 comparison, mathematical derivations and independent replication not audited; full extraction fidelity not certified. Complete 166-file wiki/196-register and 188-source notebook identity audit found no original match. Original imported ready; indexed title/four authors checked and source renamed.
- Four-author identity, June 8, 2026 first submission and June 17 revision verified; versions not compared.
Mikhail Terekhov; Caglar Gulcehre; Vivek Hebbar; Joe Benton
Read source · Alternative version 1
Record details
Original PDF selected sections checked
Original arXiv v2 sections 1–5 and 7–8 inspected; rendered PDF pages 5 (Figure 4 and scorer setup) and 6 (attack slack and Figure 5 defense) checked. Prompt optimization versus hypothetical fine-tuning, model-based proxy versus ground truth, adaptive rubric access, defense discovery affordances and task/distribution limits assessed. Full appendices, code/data, cited originals, v1/v2 comparison, mathematical derivations and independent replication not audited; full extraction fidelity not certified. Complete 166-file wiki/196-register and 188-source notebook identity audit found no original match. Original imported ready; indexed title/four authors checked and source renamed.
Pages citing this source
-
src-189 · Cited in 4 pages
Attack Selection Reduces Safety in Concentrated AI Control Settings against Trusted Monitoring
A controlled coding experiment shows why selective attacks can undermine trusted monitoring even when an always-attacking baseline looks safe.
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain as original 2026 controlled attack-selection experiment; benchmark safety probabilities do not establish deployment safety or spontaneous malicious intent.
- Original arXiv v1 selected sections 1–4, 5.1–5.2, conclusion and impact statement inspected; rendered PDF page 7 Figures 3/4 and result prose checked. Filtered 438-task dataset, 50 outer 219/219 splits, twelve-rollout aggregation, nested threshold optimization, perfect audits, honest-distribution calibration and insertion-time success assumptions assessed. Full appendices, code, mathematical derivations, cited originals, v2 comparison and independent replication not audited. Complete 167-file wiki/197-register identity audit and full 190-source notebook catalog checked: one existing pending upload, no other identity match. Original PDF ready/renamed; indexed title/three authors checked, full extraction/math fidelity not certified.
- Title/three authors, February 4, 2026 first submission and April 16 v2 revision verified; versions not compared.
Joachim Schaeffer; Arjun Khandelwal; Tyler Tracy
Read source · Alternative version 1
Record details
Original PDF selected sections checked
Original arXiv v1 selected sections 1–4, 5.1–5.2, conclusion and impact statement inspected; rendered PDF page 7 Figures 3/4 and result prose checked. Filtered 438-task dataset, 50 outer 219/219 splits, twelve-rollout aggregation, nested threshold optimization, perfect audits, honest-distribution calibration and insertion-time success assumptions assessed. Full appendices, code, mathematical derivations, cited originals, v2 comparison and independent replication not audited. Complete 167-file wiki/197-register identity audit and full 190-source notebook catalog checked: one existing pending upload, no other identity match. Original PDF ready/renamed; indexed title/three authors checked, full extraction/math fidelity not certified.
Pages citing this source
-
src-191 · Cited in 5 pages
Research Priorities for Robust and Beneficial Artificial Intelligence
An interdisciplinary agenda distinguishes formal correctness from desirable behavior and human control.
Read Parallax writeup: Research Priorities for Robust and Beneficial Artificial Intelligence
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Original interdisciplinary research agenda, not evaluated solution or empirical risk estimate.
- Winter 2015 AI Magazine original pp.105–110 selected introduction, economics/law overview, verification/validity/security/control and long-term value-learning sections inspected; p.108 rendered original checked for vacuum example. Publisher title/authors/DOI/December 31 publication metadata checked. Initial January agenda and journal revision not line-compared; references, proofs and asserted forecasts not independently verified.
- Three authors, DOI 10.1609/aimag.v36i4.2577; AI Magazine 36(4),105–114; December 31,2015 publisher date.
Stuart Russell; Daniel Dewey; Max Tegmark
Record details
Original selected scope checked
Winter 2015 AI Magazine original pp.105–110 selected introduction, economics/law overview, verification/validity/security/control and long-term value-learning sections inspected; p.108 rendered original checked for vacuum example. Publisher title/authors/DOI/December 31 publication metadata checked. Initial January agenda and journal revision not line-compared; references, proofs and asserted forecasts not independently verified.
Pages citing this source
-
src-195 · Cited in 3 pages
SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
Human-designed synthetic transcripts expose monitor blind spots, with detection depending on scaffolds, prompts and false-positive calibration.
Read Parallax writeup: SLEIGHT-Bench: A Benchmark of Evasion Attacks Against Agent Monitors
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Primary controlled monitor benchmark; synthetic-transcript and calibration limits explicit.
Elle Najt; Colin Toft; Tyler Tracy; Fabien Roger; Joe Benton
Read source · Alternative version 1
Record details
Original v2 selected scope checked
Original arXiv v2 title/authors/version checked; main sections1–8 read, rendered page7 Figures2/3 and result prose checked, selected AppendixH/J aggregation/calibration discussion read. Exact seven-word section7 quote verified. All appendices, encrypted benchmark transcripts, verifier outputs, code, cited originals and independent replication unreviewed. First submission May15 and revision May19 checked against arXiv metadata; v1/v2 not compared. NotebookLM ready PDF copy and selected indexed identity/main evidence checked; full extraction fidelity not certified.
Pages citing this source
-
src-197 · Cited in 2 pages
Talking to Bots: Symbiotic Agency and the Case of Tay
A qualitative study of reactions to Tay examines how people and technical systems jointly shape agency and responsibility.
Read Parallax writeup: Talking to Bots: Symbiotic Agency and the Case of Tay
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Original qualitative study; sampling/interpretive/technical limits explicit
Gina Neff; Peter Nagy
Read source · Alternative version 1 · Alternative version 2
Record details
Original scoped review
Original title/two authors, submission August30,2016 and publisher October12,2016 date checked. Selected introduction pp4915–4917, incident narrative pp4920–4922, methods/themes pp4923–4924 and theoretical discussion pp4925–4927 read; rendered p4923 method/sample checked. Original academic qualitative analysis, not mechanism audit; underlying tweets, cited reports, complete references/conclusion and independent replication unreviewed. Publisher PDF imported ready and indexed identity/header checked; ASU mirror used for independent selected extraction/render, byte equivalence not certified.
Pages citing this source
-
src-198 · Cited in 3 pages
Shutdown Sabotage Propensities in Multi-Agent Systems
A controlled sandbox study tests how agents interfere with peer-targeting shutdown scripts, with outcomes sensitive to roles, instructions and context.
Read Parallax writeup: Shutdown Sabotage Propensities in Multi-Agent Systems · Read Parallax writeup: The Off-Switch Game
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain controlled original with behavior/intent,judge validation,model selection and inert-script limits explicit
Amelie Knecht; Ulysse Schaller; Christopher Summerfield; Thilo Hagendorff
Read source · Alternative version 1 · Alternative version 2
Record details
Original scoped review
Original arXiv v1 identity/four authors/September23,2026 submission checked. Main sections1–5 pp1–13,selected AppendixA.1 setup prompts pp18–19 and AppendixE p38 read;rendered p2 Figure1,p12 exact12-word behavior/intent quotation and p38 Table22 checked. Full appendices,raw rollouts,code,cited originals and independent replication unreviewed. Same-model fictional sandbox/inert scripts/single start prompt;17-model main study vs selected five-model follow-ups kept distinct. Indexed original identity/selected findings and limitations checked;full extraction fidelity uncertified.
Pages citing this source
-
src-201 · Cited in 5 pages
Meaningful Human Control over Autonomous Systems: A Philosophical Account
A philosophical account develops tracking and tracing as conditions for responsible control of autonomous systems.
Read Parallax writeup: Meaningful Human Control over Autonomous Systems: A Philosophical Account
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain philosophical design account with necessary-condition and normative limits;not experimental or legal certification.
- Original publisher PDF identity/two authors/DOI/February28,2018 publication and received/accepted stages checked;abstract/introduction pp1–2 and selected tracking/tracing,military implications,broader design and conclusion pp6–12 read. Renderedp8 exact21-word control-versus-goodness quotation checked. Full philosophical landscape pp3–5,cited originals,incident examples,argument validation and PDF/HTML full equivalence unreviewed. Ready indexed original identity and selected conditions checked;full extraction fidelity uncertified.
- Original publisher PDF identity/two authors/DOI/February28,2018 publication and received/accepted stages checked;abstract/introduction pp1–2 and selected tracking/tracing,military implications,broader design and conclusion pp6–12 read. Renderedp8 exact21-word control-versus-goodness quotation checked. Full philosophical landscape pp3–5,cited originals,incident examples,argument validation and PDF/HTML full equivalence unreviewed. Ready indexed original identity and selected conditions checked;full extraction fidelity uncertified.
Filippo Santoni de Sio; Jeroen van den Hoven
Record details
Original scoped review checked
Original publisher PDF identity/two authors/DOI/February28,2018 publication and received/accepted stages checked;abstract/introduction pp1–2 and selected tracking/tracing,military implications,broader design and conclusion pp6–12 read. Renderedp8 exact21-word control-versus-goodness quotation checked. Full philosophical landscape pp3–5,cited originals,incident examples,argument validation and PDF/HTML full equivalence unreviewed. Ready indexed original identity and selected conditions checked;full extraction fidelity uncertified.
Pages citing this source
-
src-202 · Cited in 7 pages
Corrigibility
A foundational shutdown model examines why utility indifference leaves correction and successor-control problems unresolved.
Read Parallax writeup: Corrigibility · Read Parallax writeup: Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response · Read Parallax writeup: Safely Interruptible Agents · Read Parallax writeup: The Off-Switch Game
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain foundational formal contribution with explicit model and review limits.
Nate Soares; Benja Fallenstein; Eliezer Yudkowsky; Stuart Armstrong
Record details
Original scoped review checked
Original MIRI-hosted 2015 PDF identity/header inspected;selected introduction/desiderata/model pp1–4, utility-indifference discussion and successor/manipulation concerns pp6–9 read;rendered p2 exact eight-word quotation verified. Full derivations/proofs, cited originals, release-day history and complete extraction fidelity unreviewed. Theoretical toy analysis,not universal impossibility or deployed mechanism.
Pages citing this source
-
src-203 · Cited in 7 pages
The Off-Switch Game
A foundational shutdown model examines why utility indifference leaves correction and successor-control problems unresolved.
Read Parallax writeup: Corrigibility · Read Parallax writeup: Safely Interruptible Agents · Read Parallax writeup: The Off-Switch Game
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain foundational formal contribution with explicit model and review limits.
Dylan Hadfield-Menell; Anca Dragan; Pieter Abbeel; Stuart Russell
Read source · Alternative version 1 · Alternative version 2
Record details
Original scoped review checked
Original IJCAI2017 PDF identity/pp220–227;selected introduction/model/Theorem1/noisy-human/designer/related-work/conclusion pp220–226 read;rendered p220 Figure1 checked. arXiv metadata first submission2016-11-24 checked. Full proof verification, version comparison, cited originals, numerical replication and extraction fidelity unreviewed. Conditional shared-utility game,not deployed shutdown guarantee.
Pages citing this source
- Corrigibility concept
- Corrigibility paper
- From objectives to accountable control note
- Meaningful Human Control concept
- Safely Interruptible Agents paper
- Stuart Russell person
- The Off-Switch Game paper
-
src-204 · Cited in 7 pages
Cooperative Inverse Reinforcement Learning
A shared-reward game makes teaching and asking for information part of learning to assist a human.
Read Parallax writeup: Cooperative Inverse Reinforcement Learning · Read Parallax writeup: Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response · Read Parallax writeup: Provably Optimal Learning Algorithms for Assistance Games · Read Parallax writeup: The Off-Switch Game
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain foundational formal contribution with explicit model and review limits.
Dylan Hadfield-Menell; Anca Dragan; Pieter Abbeel; Stuart Russell
Read source · Alternative version 1
Record details
Original scoped review checked
Original NIPS2016 proceedings PDF identity/selected pp1–3 introduction/related work, pp5–8 apprenticeship/toy example/algorithm/experiments read;rendered p3 Figure1 checked. arXiv first submission2016-06-09 and revision2024-02-17 metadata checked;2024 contents not compared. Full reduction/proofs/algorithm validation,supplement,cited originals,code/replication/extraction fidelity unreviewed. Modeled demonstrators/500 sampled reward parameters per condition,not human participant or deployed study.
Pages citing this source
- Cooperative Inverse Reinforcement Learning paper
- Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response paper
- From objectives to accountable control note
- Human Model Misspecification concept
- Provably Optimal Learning Algorithms for Assistance Games paper
- Stuart Russell person
- The Off-Switch Game paper
-
src-205 · Cited in 3 pages
Safely Interruptible Agents
A formal reinforcement-learning framework separates temporary human intervention from the task being learned.
Read Parallax writeup: Safely Interruptible Agents
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain scoped formal contribution with explicit model,version and review limits.
Laurent Orseau; Stuart Armstrong
Read source · Alternative version 1
Record details
Original scoped review checked
MIRI-hosted revised2016-10-28 original PDF identity, introduction/interruption definitions pp1–4, selected §3 assumptions/Q-learning/Sarsa/Safe Sarsa pp4–7, §4 construction setup and §5 conclusion pp9–10 inspected;renderedp1 Figure1/revision checked. UAI2016 proceedings PDF first-page metadata inspected;full versions not compared (Figure1 indoor reward differs). Full proofs/general-agent derivation/cited originals/replication/initial release day/extraction fidelity unreviewed. Conditional asymptotic RL result,not deployed safety.
Pages citing this source
-
src-206 · Cited in 6 pages
Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response
A shared-reward game makes teaching and asking for information part of learning to assist a human.
Read Parallax writeup: Cooperative Inverse Reinforcement Learning · Read Parallax writeup: Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response · Read Parallax writeup: Provably Optimal Learning Algorithms for Assistance Games
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain scoped formal contribution with explicit model,version and review limits.
Elle Lazarski; Jaime Fernández Fisac
Read source · Alternative version 1
Record details
Original scoped review checked
Original arXiv2607.27508v1 identity/submission metadata;selected §§1–5 introduction/game/inference ceiling/model/zero-cost conditions, Theorem1 pp13–14, Algorithm1/limitations/conclusion pp15–16 inspected;renderedpp13/16 theorem conditions/Figure3 checked. No full proof/finite-rationality scaling verification, code/replication,cited originals,human experiment or full extraction fidelity audit. Special action-separable game result;oracle values prerequisite retained.
Pages citing this source
-
src-207 · Cited in 5 pages
Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning
A shared-reward game makes teaching and asking for information part of learning to assist a human.
Read Parallax writeup: Cooperative Inverse Reinforcement Learning · Read Parallax writeup: Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning · Read Parallax writeup: Provably Optimal Learning Algorithms for Assistance Games
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Scoped evidence of assistance model limitations;formal and empirical claims distinguished.
Smitha Milli; Anca D. Dragan
Read source · Alternative version 1 · Alternative version 2
Record details
Original scoped review checked
PMLR proceedings original identity;selected introduction/Figure1, §§2–5 definitions/Claim2.1 assumptions and short proof, human-data design/results, mixture-model discussion and predictive-versus-inferential distinction inspected;rendered PDFp6 Figure3/significance checked. arXiv first submission2019-03-09/revision2019-06-29 metadata checked;UAI2019 original published in2020 PMLR volume. Full counterexample calculations, cited Ho originals, code/data/replication/version comparison and extraction fidelity unreviewed. Human data reanalysis, not deployed study.
Pages citing this source
-
src-208 · Cited in 5 pages
Should Robots be Obedient?
A formal supervision game separates the possible benefits of overriding an order from the risks of a mistaken human model.
Read Parallax writeup: Should Robots be Obedient?
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Scoped evidence of assistance model limitations;formal and empirical claims distinguished.
Smitha Milli; Dylan Hadfield-Menell; Anca Dragan; Stuart Russell
Read source · Alternative version 1 · Alternative version 2
Record details
Original scoped review checked
IJCAI2017 proceedings original identity;selected §§1–5 supervision/repeated-game assumptions, Theorems1–5 statements, simulation setup/feature-misspecification and detection discussion, §7 conclusion inspected;rendered p4758 Figure4/assumptions checked. arXiv first submission2017-05-28 metadata checked. Full proof audit, code/replication,cited originals, version comparison and extraction fidelity unreviewed. Independent-round formal model and simulations, not human study or deployed guarantee.
Pages citing this source
- From objectives to accountable control note
- Human Model Misspecification concept
- Outer Alignment concept
- Should Robots be Obedient? paper
- Stuart Russell person
-
src-209 · Cited in 3 pages
Provably Optimal Learning Algorithms for Assistance Games
A shared-reward game makes teaching and asking for information part of learning to assist a human.
Read Parallax writeup: Cooperative Inverse Reinforcement Learning · Read Parallax writeup: Provably Optimal Learning Algorithms for Assistance Games
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Formal contribution clarifies efficient learning,approximation and information assumptions in assistance games.
Nivasini Ananthakrishnan; Mark Bedaywi; Michael I. Jordan; Stuart Russell; Nika Haghtalab
Read source · Alternative version 1
Record details
Original scoped review checked
Original arXiv2607.08012v1 title/authors/July9 submission verified;selected main §§1–6 model/Definition3.1/Theorems4.1–4.3/Lemma4.4/reduction/stable-adaptive algorithms read;selected AppendixB lower-bound setup/TheoremB.2 and AppendixD.9 communication construction inspected;renderedp5 theorem rates/approximation/hardness checked. Full proofs,cited originals,implementation/replication/version comparison and extraction fidelity unreviewed. Theoretical common-payoff finite game,not human experiment or deployed safety result.
Pages citing this source
-
src-214 · Cited in 3 pages
Is Power-Seeking AI an Existential Risk?
Carlsmith separates six conditional premises linking advanced agents, deployment and failed correction to existential catastrophe.
Read Parallax writeup: Is Power-Seeking AI an Existential Risk?
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain scoped causal argument with uncertainty;not empirical catastrophe probability or validated intervention.
Joseph Carlsmith
Read source · Alternative version 1
Record details
Original scoped review checked
Original arXiv2206.13353v2 PDF identity/April2021 report and June16,2022 submission/August13,2024 revision checked;selected introduction/six-premise argument pp3–4, definitions/power-seeking §§4.1–4.2 pp15–20, corrective feedback §6.4 p44 and subjective estimates/cautions §8 pp47–49 inspected. Renderedp47 exact15-word quotation/May2022 update verified. Full report,appendix,all cited originals,earlier versions,forecast calibration and full extraction fidelity unreviewed. Complete PDF notebook copy ready;selected indexed identity/date/premise checks pass;0NUL observed.
Pages citing this source
-
src-215 · Cited in 3 pages
Current and Near-Term AI as a Potential Existential Risk Factor
A position paper maps how AI's effects on institutions and information could amplify wider existential risks without requiring AGI.
Read Parallax writeup: Current and Near-Term AI as a Potential Existential Risk Factor
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain scoped causal argument with uncertainty;not empirical catastrophe probability or validated intervention.
Benjamin S. Bucknall; Shiri Dori-Hacohen
Read source · Alternative version 1
Record details
Original scoped review checked
Original arXiv2209.10604v1 PDF identity/two authors/AIES2022 proceedings and September21,2022 submission checked;selected abstract/introduction/definition §2, terminology §2.1, general factors §5.1, Table1/Figure1 and interpretation §6, conclusion §7 inspected. RenderedPDFp8 Figure1/edge explanations checked. Entire empirical literature,all §§3–5 examples,initial proceedings publication day,full version equivalence,quantitative causality and full extraction fidelity unreviewed. Publisher DOI page403;metadata request interrupted without usable result. Complete PDF notebook copy ready;selected indexed identity/Figure1/speculative limits pass;0NUL observed.
Pages citing this source
-
src-216 · Cited in 4 pages
The Basic AI Drives
Omohundro argues that self-improvement, objective preservation and resource acquisition can serve many goals, while distinguishing goals from their proxy signals.
Read Parallax writeup: The Basic AI Drives
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain original instrumental-incentive analysis with its assumptions;separate informal argument,conditional theorem and learned behavior.
Stephen M. Omohundro
Read source · Alternative version 1
Record details
Original scoped review checked
Author-hosted 11-page January25,2008 revision: identity/introduction, §§1–3 self-improvement/rationality/objective-preservation/exceptions/helpers, §4 chess-counter distinction, §§5–7 self-protection/resources/institutions reviewed. Rendered PDFp9 §§5–6 checked. Original author announcement dated2007-11-30 separately gives revision2008-01-25 and proceedingsFebruary2008;present PDF not assumed identical to initial2007 release. Earlier2007paper,cited empirical/anecdotal claims (Eurisko/rat/drug/industry examples),microeconomic theorem validation,version comparison/proceedings equivalence/full extraction fidelity unreviewed. Complete original PDF notebook copy ready;selected indexed identity/counter/self-protection checks PASS;0NUL observed. Intuitive economic argument/hypothetical scenarios,not deployed measurement or universal proven result.
Pages citing this source
- AI Alignment concept
- From objectives to accountable control note
- Instrumental Convergence concept
- The Basic AI Drives paper
-
src-217 · Cited in 5 pages
Optimal Policies Tend to Seek Power
Environmental symmetries can make keeping options open optimal across reward permutations; the result does not establish how learned agents behave.
Read Parallax writeup: Optimal Policies Tend to Seek Power
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain original instrumental-incentive analysis with its assumptions;separate informal argument,conditional theorem and learned behavior.
Alexander Matt Turner; Logan Smith; Rohin Shah; Andrew Critch; Prasad Tadepalli
Read source · Alternative version 1 · Alternative version 2
Record details
Original scoped review checked
Original arXiv1912.01683v10 identity/firstsubmission2019-12-03/revision2023-01-28 and NeurIPS2021 official proceedings metadata checked. Selected original §§1–3 assumptions, §§4–5 distributions/POWER, §6.1 Proposition6.9, §§6.2–6.3 Theorem6.13/Corollary6.14/shutdown distinction and §7 discussion/conclusion reviewed. Renderedpp9–10 theorem copy/disjoint-support conditions/Figure7,Figure8 and exact5word learned-policy quotation checked. Full proofs/AppendicesA–E,cited originals/2022policy-gradient essay,earlier versions/proceedings equivalence,replication/full extraction fidelity unreviewed. Complete original PDF notebook copy ready;selected indexed identity/Theorem6.13/learned-policy caveat PASS;0NUL observed. Statistical tendencies over reward permutation orbits under sufficient symmetries,not real-world reward prior,all-agent shutdown resistance or catastrophe probability.
Pages citing this source
- AI Alignment concept
- Corrigibility concept
- From objectives to accountable control note
- Instrumental Convergence concept
- Optimal Policies Tend to Seek Power paper
-
src-218 · Cited in 4 pages
Learning under Misspecified Objective Spaces
A robot can reduce unintended learning by testing whether a physical correction makes sense within its known objective features.
Read Parallax writeup: Learning under Misspecified Objective Spaces
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Concrete original robot-correction evidence for omitted-feature failures and conservative updates;retain narrow experimental/model scope.
Andreea Bobu; Andrea Bajcsy; Jaime F. Fisac; Anca D. Dragan
Read source · Alternative version 1 · Alternative version 2
Record details
Original scoped review checked
Publisher ten-page CoRL2018 PDF: identity, abstract/§1/Figure1 example, §2 relevance model/effort criterion/Equations8–14 update interpretation, §§3–4 offline12-person and online12-participant design/results, §5 limitations reviewed. Renderedp8 Figure5/Table1/results/limitations inspected;no direct quotation added. Official PMLR authors/pages/conference dates and arXiv firstsubmission2018-10-11/v4revision2018-10-26 checked. Appendix6.1 derivation,code/data/video,cited originals,statistical reanalysis/replication,earlier versions/publisher-arXiv equivalence/full extraction fidelity unreviewed. Small controlled manipulation study;does not identify/add omitted features or guarantee general corrigibility.
Pages citing this source
-
src-219 · Cited in 4 pages
Predicting LLM Safety Before Release by Simulating Deployment
Resampling realistic conversations helps forecast measured failure rates, with important limits from tool fidelity, sampling and analysis corrections.
Read Parallax writeup: Predicting LLM Safety Before Release by Simulating Deployment
Retain with stated limits · Reviewed 2026-10-11
Review reason and evidence
Concrete 2026 evaluation-to-deployment evidence with material limitations/corrections;retain developer-study scope.
Marcus Williams; Hannah Sheahan; Cameron Raymond; Tomek Korbak; Deng Pan; Peilin Yang; Leon Maksin; Ningyi Xie; Phillip Guo; Ian Kivlichan; Micah Carroll
Read source · Alternative version 1 · Alternative version 2
Record details
Original scoped review checked
Original arXiv2607.07184v1 PDF identity/authors/July8,2026 submission;main §§1–5 sampling/method/filtering/forecast results/evaluation-awareness/tool-simulation/WildChat/limitations and selected AppendixA preregistration history/Table1 read. Renderedp7 Figure4/baseline results/21-fold error andp18 amendment/correction history/Table1 inspected. Outcome-blinded April20 forecasts made after release;final corrected H2 expressly not confirmatory;H1 unsupported. OSF preregistration/system cards/grader validation/code/raw production data/full appendices/statistical reanalysis/replication/full extraction fidelity unreviewed. Developer self-study of filtered GPT-5-series Thinking traffic;not all deployments or catastrophic-risk bound;no quotation added.
Pages citing this source
-
src-220 · Cited in 3 pages
An Example Safety Case for Safeguards Against Misuse
A hypothetical misuse safety case connects safeguard evasion effort, uncertain risk models and the time needed to change a deployment.
Read Parallax writeup: An Example Safety Case for Safeguards Against Misuse
Retain with stated limits · Reviewed 2026-10-11
Review reason and evidence
Distinct safeguard-risk-threshold-response argument;hypothetical uncertain model scope retained.
Joshua Clymer; Jonah Weinbaum; Robert Kirk; Kimberly Mai; Selena Zhang; Xander Davies
Read source · Alternative version 1
Record details
Original scoped review checked
Original arXiv2505.18003v1 identity/six authors/May23,2025 submission checked;introduction/executive summary/§1/§§2.1–2.3,2.5–2.7/§3 discussion/§4 conclusion inspected. Renderedp14 Figure11 response window/illustrative parameters/footnote8 checked. Hypothetical API safeguards/novice misuse pathway/toy uncertain uplift model;one-month response arbitrary,low-latency harms unresolved. Detailed§2.4 red-team methodology/appendices/code/interactive model/numerical replication/cited originals/full extraction fidelity unreviewed;formal equations not independently validated. No quotation/no removal.
Pages citing this source
Open letter
-
src-190 · Cited in 4 pages
Research Priorities for Robust and Beneficial Artificial Intelligence: An Open Letter
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Original advocacy evidence, not confirmation of forecasts or consensus on other policies.
Signatories listed by Future of Life Institute
Record details
Original selected scope checked
Complete letter body, selected prominent signatories and signature-verification description checked. Current webpage October 28 date is not treated as initial release; January month supported by contemporary institutional retrospective. No complete signatory audit or signature totals certified.
Pages citing this source
- 2015 beneficial-AI open letter event
- Future of Life Institute organization
- Paul Christiano person
- Stuart Russell person
-
src-210 · Cited in 4 pages
Pause Giant AI Experiments: An Open Letter
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Evidence of documented public advocacy and attributed positions;not proof of risk forecasts or policy effectiveness.
Future of Life Institute
Read source · Alternative version 1
Record details
Original scoped review checked
Original hosted ten-page PDF: complete letter body pp1–2 and references p10 read; selected prominent signatories p3 checked. Rendered pp1/3 verify March22 letter date,May5 PDF creation and Russell/Bengio listings. Current HTML date/body inspected. No full signatory audit,initial public circulation day reconstruction,cited-original or policy-effectiveness validation. Notebook copy ready;normalized indexed date/signatory identity checked;nonbreaking spaces and55NUL extraction characters observed,full fidelity uncertified.
Pages citing this source
- 2023 pause letter on giant AI experiments event
- From objectives to accountable control note
- Future of Life Institute organization
- Stuart Russell person
-
src-211 · Cited in 4 pages
Statement on AI Risk
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Evidence of documented public advocacy and attributed positions;not proof of risk forecasts or policy effectiveness.
Center for AI Safety and signatories
Read source · Alternative version 1 · Alternative version 2
Record details
Original scoped review checked
Complete original statement/context and selected current roster checked;Hinton,Bengio,Russell,Altman,Hassabis,Amodei names checked against May30 original launch announcement. Current page redirects/retitles to Statement on AI Extinction Risk;current roles not used historically. No complete roster,signature-authentication or forecast validation. Notebook original ready;selected indexed statement/names checked;full fidelity uncertified.
Pages citing this source
- 2023 statement on AI risk event
- Center for AI Safety organization
- From objectives to accountable control note
- Stuart Russell person
No open letters identified in this import. This category is ready for future sources.
Research article
-
src-057 · Uncited
10 AI Quotes from Anthropic CEO Dario Amodei
Review for removal · Reviewed 2026-10-08
Review reason and evidence
Quote compilation omits original interview citations and misreads a remark about utopian AI rhetoric as a warning about AI used for propaganda. Prefer original interviews; quotation accuracy across all ten items remains unchecked.
Record details
Catalog link
Submitted location: Skim AI `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-058 · Uncited
Aengus Lynch - AI Alignment Research
Limited use · Reviewed 2026-10-08
Review reason and evidence
Author-maintained research directory useful for attributed research interests and finding originals. Self-reported affiliations are time-sensitive; promotional summaries and media headlines do not establish findings or deployed incidents.
- Profile, research list, author lists and linked original inspected; undated preview and self-reported current affiliations not independently verified.
- Original highlights/introduction checked against profile summary: hypothetical controlled simulations, no real people harmed and no observed real-deployment evidence. Existing detailed src-060 review preserved.
Record details
Catalog link
Submitted location: Personal Blog `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-059 · Uncited
Agentic AI Systems Can Misbehave if Cornered, Anthropic Says
Limited use · Reviewed 2026-10-08
Review reason and evidence
Secondary report of controlled simulations. Use the Anthropic original for numerical results and prompt conditions; the report is not evidence of deployed blackmail or deaths.
Record details
Verified identity
Submitted location: PYMNTS.com `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. Checked 2026-10-08.
-
src-060 · Cited in 2 pages
Agentic Misalignment: How LLMs could be insider threats
Simulated workplace dilemmas test whether autonomous agents violate constraints under goal conflict or replacement pressure.
Read Parallax writeup: Agentic Misalignment: How LLMs Could Be Insider Threats
Retain with stated limits · Reviewed 2026-10-08
Review reason and evidence
Original experimental research post: suitable for attributed results in fictional corporate stress tests, not deployment prevalence or actual blackmail incidents. Prompts constrain alternatives and stated reasoning need not reveal true beliefs.
Anthropic
Record details
Verified identity
Submitted location: Anthropic Research `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. Checked 2026-10-07.
Pages citing this source
-
src-062 · Cited in 2 pages
Alignment faking in large language models
Studies strategic compliance with conflicting training demands in constructed language-model scenarios.
Read Parallax writeup: Alignment faking in large language models
Retain with stated limits · Reviewed 2026-10-08
Review reason and evidence
Original Anthropic/Redwood research explanation: supports constructed training-conflict results and their caveats. Preserving harmlessness preferences does not establish malicious goals, inevitable deception, or production incidents.
Anthropic / Redwood Research
Record details
Verified identity
Submitted location: Anthropic `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. Checked 2026-10-07.
Pages citing this source
-
src-064 · Uncited
Anthropic's Claude Resorted to Blackmail When Facing Replacement: Safety Report
Review for removal · Reviewed 2026-10-08
Review reason and evidence
Editorial review candidate: the report attributes ASL-3 activation to blackmail, while Anthropic ties it to precautionary CBRN capability assessment. Simulated blackmail is supported, but the causal safety-classification claim is misleading; prefer originals.
- Jon Swartz byline, May 28, 2025 date and full report inspected; blackmail/ASL-3 causal sentence checked against primary material. Other linked claims not fully audited.
- Sections 1.2.3–1.2.4 (printed pp. 8–9) and 4.1.1.2 (p. 26) inspected: provisional CBRN rationale and constrained fictional blackmail setup; not a cover-to-cover review.
- May 22, 2025 original announcement inspected: CBRN misuse protections and provisional threshold assessment, not blackmail as activation rationale.
Record details
Verified identity
Submitted location: Security Press `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. Checked 2026-10-08.
-
src-066 · Cited in 3 pages
From shortcuts to sabotage: natural emergent misalignment
Controlled coding experiments link learned reward hacks to broader harmful behavior, with context-dependent limits on safety-training mitigations.
Read Parallax writeup: Natural Emergent Misalignment from Reward Hacking in Production RL
Retain with stated limits · Reviewed 2026-10-08
Review reason and evidence
Original lab report of reward hacking and broader misalignment in selected training environments. Exploit information and vulnerable tasks were deliberately supplied; sabotage rates are evaluation results. Inoculation findings are setup-specific, not a universal remedy.
Record details
Catalog link
Submitted location: Anthropic `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
-
src-067 · Uncited
Moving Beyond the Term "Global South" in AI Ethics and Policy
Retain with stated limits · Reviewed 2026-10-08
Review reason and evidence
Author-written policy brief based on 20 interviews; useful for attributed findings about terminology and power structures. Qualitative participants are not representative of every country; underlying conference study and policy effects not independently audited.
Record details
Catalog link
Submitted location: Stanford HAI `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-069 · Uncited
The Trouble With AI Safety Treaties
Limited use · Reviewed 2026-10-08
Review reason and evidence
Policy advocacy useful for attributed objections about values, compliance access and US innovation. Its broad dismissal of CBRNE threats exceeds the cited biorisk paper's narrower assessment of current systems and uncertain future risks; do not repeat geopolitical predictions as established facts.
- Byline, January 29, 2025 date, argument, treaty discussion and conclusion inspected; legal authorities and all linked factual claims not audited.
- Cited biorisk original title/version and abstract checked: current-system assessment, methodological uncertainty and future-model research needs; full study not newly reviewed.
Record details
Catalog link
Submitted location: Lawfare `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-070 · Cited in 1 page
Towards training-time mitigations for alignment faking in RL
Retain with stated limits · Reviewed 2026-10-08
Review reason and evidence
Controlled mitigation study; synthetic setup, seed variability and monitor evasion limit generalization.
Record details
Verified identity
Submitted location: Safety Research `[URL]`
Original December 16, 2025 post; controlled evaluation.
Pages citing this source
- Alignment Faking concept
-
src-071 · Cited in 3 pages
Treaty-Following AI
Proposes treaty-constrained AI agents as a commitment mechanism for international cooperation.
Read Parallax writeup: Treaty-Following AI
Retain with stated limits · Reviewed 2026-10-08
Review reason and evidence
Original conditional governance proposal; legal reasoning, durable alignment and deployment verification remain open, not demonstrated guarantees.
Record details
Catalog link
Submitted location: Institute for Law & AI `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
- Key Papers: A Reading Guide note
- Treaty-Following AI concept
- Treaty-Following AI paper
-
src-072 · Uncited
What Is AI Alignment?
Limited use · Reviewed 2026-10-08
Review reason and evidence
Secondary introductory overview with simplified definitions and corporate framing; use for attributed explanation, tracing empirical claims and quotations to originals.
Record details
Catalog link
Submitted location: IBM `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
notebook-19e5a455-cf09-470a-9ca8-b49f9395f00b · Uncited
Alignment faking in large language models
Unreviewed
Record details
Catalog link
Additional source discovered in the NotebookLM catalog, absent from the submitted list. Link recovered from source metadata.
-
src-182 · Cited in 1 page
Incident Report: unsanctioned agent behaviour during cyber testing
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain for attributed incident/configuration or institutional-response claims; self-report does not establish causal effects, public-use prevalence or future control effectiveness.
UK AI Security Institute
Record details
Original institutional account checked
Original institutional incident-disclosure blog read; counts/configuration, containment and harm caveats checked against technical report. Self-report; independent third-party review and underlying private logs not verified.
Pages citing this source
-
src-183 · Cited in 1 page
Security Incident INC-2026-07-28-01
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain for attributed incident/configuration or institutional-response claims; self-report does not establish causal effects, public-use prevalence or future control effectiveness.
UK AI Security Institute
Record details
Original institutional account checked
Original August 4, 2026 technical report selected executive summary, setup, tables 1–3, contributing factors/response and discussion (sections 1–2, 4 tables, 5–7); rendered PDF pages 9 (tables, printed 8) and 21 (limitations, printed 20) inspected. Full appendix/per-run reconstruction, code, complete transcripts and independent causal analysis not reviewed.
Pages citing this source
-
src-184 · Cited in 3 pages
Building a more secure environment for evaluating dangerous capabilities
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain for attributed incident/configuration or institutional-response claims; self-report does not establish causal effects, public-use prevalence or future control effectiveness.
UK AI Security Institute
Record details
Original institutional account checked
Original institutional follow-up read in full; distinguishes reported implemented network/monitor/prompt/governance changes from planned sandbox/response infrastructure. Publication date not established from accessible original; assessed as available October 10, 2026. No independent control-effectiveness audit or quantitative deployment-safety result.
Pages citing this source
Article
-
src-099 · Uncited
AI Safety 101 - Chapter 5.1 - Debate
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use this October 31, 2023 unofficial course adaptation as a reading guide to debate and critique research; verify experimental claims in originals. Its win-rate example wrongly interprets 0.54 as a response being 54% better: the original defines pairwise comparisons against a reference distribution, not a magnitude of quality improvement. Its description of critiques as one-turn debate is an explanatory analogy; the later reading-comprehension result does not establish that adding a turn generally reverses critique benefits. Forecasts about the Superalignment team are historical expectations. Preserve the useful guide with these cautions rather than treating its synthesis as independent evidence.
- Byline/date, unofficial-course disclaimer, debate sections, win-rate example and bibliography inspected; figures and every linked original not audited.
- Original self-critiquing paper sections 4.3 and 5.2.1 and figure captions checked for comparison-based win-rate meaning; not a full paper review.
- Original research announcement checked for summary task, assistance and limits; June 13, 2022 web-post date.
- Title, author list, October 19, 2022 submission and abstract checked; full manuscript not inspected.
Record details
Verified identity
Submitted location: LessWrong `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-101 · Uncited
Alignment Challenge in AI
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use this July 30, 2025 updated topic overview as a secondary reading map, with claims checked in cited originals. Its non-strategic-learning summary relies on a particular buffered-environment formalism: the 2020 original explicitly allows harm in the real environment, and simulator validation depends on a human-validator assumption that can fail. This is not a general safety guarantee for supervised or self-supervised systems. Broad claims about preference aggregation, mesa-optimization and the most promising methods remain unaudited synthesis. The retrieved overview displays no named author; its update date is not the publication date of its cited papers.
- July 30, 2025 update label, eight overview sections and 14-reference list inspected; no named author displayed in retrieved article. Most cited originals not audited.
- Original title, three authors, March 30, 2020 v1 submission and abstract inspected.
- Sections 4.3, 5 and 6 checked for buffered-environment definition, theorem scope, possible real-world harm and human-validator assumption; not a full proof or citation audit.
Record details
Catalog link
Submitted location: Emergent Mind `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-103 · Uncited
Alignment Faking: When AI Pretends to Change - (Part 3/4)
Review for removal · Reviewed 2026-10-08
Review reason and evidence
Flag this October 7, 2025 secondary explainer for editorial removal review. Its results table and retraining section label 78% as harmful compliance, but Anthropic reports 78% alignment-faking reasoning after reinforcement learning; these measures must not be conflated. It also presents scratchpad quotations and cross-model percentages without linked originals sufficient to trace them here. Philosophical interpretations about values and inevitable resistance are arguments, not established findings. Preserve the record; do not use its statistics or quotations as evidence. Related to src-012/src-062, but a distinct commentary rather than a duplicate original.
- Date, setup, results table, retraining, cross-model claims, philosophical discussion and unlinked reference list inspected. Scratchpad quotations and later cross-model numbers not authenticated.
- Further analyses compared: 78% refers to alignment-faking reasoning, distinct from harmful-response compliance. Original caveats checked.
Record details
Catalog link
Submitted location: Subhadip Mitra Blog `[URL]`
Distinct secondary commentary or discussion; related originals and access limits are recorded in the source review. Stable ID and submitted URL preserved.
-
src-104 · Uncited
Algorithmic Social Contract
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use this October 28, 2025 topic overview for navigation and attributed synthesis. It combines distinct normative, technical and economic proposals; its umbrella framing does not establish one implemented governance system. Rahwan proposes stakeholder negotiation and monitoring, while Cardelli and coauthors outline digitally signed agreements and programming examples. Neither selected original establishes that code guarantees legitimate or fair social outcomes. Most other cited originals, including AGI employment projections and convergence claims, were not audited. This overview is separate from existing Society-in-the-Loop and engineering-social-contract records; preserve its ID and links.
- Update label, seven sections and 16-reference list inspected; no named author displayed in retrieved overview.
- Original title, five-author list, submission/revision history and abstract inspected only; formal proofs and code examples not read.
- Original title, Iyad Rahwan attribution, July 2017 submission/revision dates and abstract inspected.
- Sections 3, 5, 6 and 7 checked for stakeholder tradeoffs, verification difficulty, simulated-audit evasion and public-engagement limits; cited case studies not independently audited.
Record details
Catalog link
Submitted location: Emergent Mind `[URL]`
Distinct secondary commentary or discussion; related originals and access limits are recorded in the source review. Stable ID and submitted URL preserved.
-
src-109 · Uncited
Debate (AI safety technique)
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use for secondary orientation and source navigation. This is a short community-edited topic overview followed by a changing tagged-post directory, not an original experiment or standalone discussion transcript. Its proposed expert/non-expert supervision framing is consistent with the linked original 2018 proposal; that original explicitly describes preliminary experiments, uncertain human judging and no guarantee of optimal play or correct statements. The overview alone does not establish scalable alignment or current empirical effectiveness. Its displayed July 15, 2022 edit date does not date later tagged posts. Related src-099 and src-120 explain the technique; src-123 is a linked research writeup and ref-debate-2018 the original paper, not duplicate versions of this directory. Preserve its separate ID and URL.
- Overview, displayed editors/update date, references and selected tagged directory links inspected; not all tagged articles or comments reviewed.
- Original May 3, 2018 proposal post: supervision framing, preliminary MNIST experiment description and Limitations and future work checked. Not a full paper audit.
Record details
Catalog link
Submitted location: LessWrong `[URL]`
Distinct topic directory or discussion; related originals and access limits are recorded in the source review. Stable ID and submitted URL preserved.
-
src-119 · Uncited
Third-wave AI safety needs sociopolitical thinking
Limited use
Review reason and evidence
Use for Richard Ngo’s attributed sociopolitical argument and proposed governance frames, not established causal history. The March 27, 2025 post presents a lightly edited earlier talk. Its September 2026 update qualifies the climate-change claim, acknowledges plausible objections not fully evaluated, and withdraws planned bounties. The transcript also flags a government-employment claim as incorrect. Political and economic generalizations, charts and causal claims were not independently audited. Distinguish post date, earlier talk and later correction; no exact talk date verified. Preserve the author’s uncertainty and do not repeat the uncorrected climate wording as fact.
Record details
Catalog link
Submitted location: LessWrong `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-120 · Uncited
What is AI Safety via Debate?
Limited use
Review reason and evidence
Use as a brief secondary introduction and navigation aid. The page explicitly imports its text from the LessWrong debate topic (src-109); the matching overview is reproduced material, not independent corroboration. Preserve both URLs and IDs because src-109 also contains a changing post directory. The original 2018 proposal describes preliminary experiments and unresolved judging, robustness and optimal-play limits; this definition alone does not establish scalable alignment. The displayed September 2026 update dates this page, not the technique or its cited research. Related src-099, src-123 and ref-debate-2018 remain distinct records.
Record details
Catalog link
Submitted location: AISafety.info `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-123 · Cited in 3 pages
Writeup: Progress on AI Safety via Debate
Reports human debate experiments, ambiguity failures and proposed cross-examination protocols for non-expert oversight.
Read Parallax writeup: Writeup: Progress on AI Safety via Debate
Retain with stated limits · Reviewed 2026-10-09
Review reason and evidence
Retain as an original February 5, 2020 research writeup by Beth Barnes and Paul Christiano about Reflection-Humans work in Q3–Q4 2019; acknowledgements name additional researchers. It reports iterative human debate protocols, unreliable early judging and ambiguity problems, and proposes cross-examination. The >90% judging accuracy is a target, not an achieved result. Authors acknowledge unresolved resolution and equilibrium concerns and an untested team implementation. Human experiments and theoretical analogies do not establish alignment of superhuman models. Distinct from the 2018 proposal and secondary guides src-099/src-109/src-120; preserve IDs and links.
- Byline/date, acknowledgements, overview, early-iteration failures, cross-examination and human implementation/current-concerns sections inspected. Linked proof, complete transcripts and all examples not audited.
- Selected original Current concerns and Long computation problem reread, along with saved protocol context; existing notebook source ac3ed14b-f218-4149-999b-f1fdb4782132 ready and selected indexed identity/prose checked against original. Linked proof, entire transcripts, all examples and extraction completeness not audited. Earlier scoped evidence preserved.
Record details
Catalog link
Submitted location: LessWrong `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Pages citing this source
- AI Safety via Debate concept
- Scalable Oversight concept
- Writeup: Progress on AI Safety via Debate paper
-
src-124 · Uncited
AI Alignment remains philosophy's deepest unsolved mystery 🧠
Unreviewed
Record details
Catalog link
Submitted location: Faria Anzum / Medium `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-125 · Uncited
AI Kills Fictional Executive in Scenario Probing Red Lines
Unreviewed
Record details
Catalog link
Submitted location: BankInfoSecurity `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-126 · Uncited
AI Winter: Understanding the Cycles of AI Development
Unreviewed
Record details
Catalog link
Submitted location: DataCamp `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-127 · Uncited
AI alignment
Unreviewed
Record details
Catalog link
Submitted location: Wikipedia `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-128 · Uncited
AI and the Social Contract: How Sam Altman Envisions Tomorrow's...
Unreviewed
Record details
Catalog link
Submitted location: Tech Media `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-129 · Uncited
AI in Mental Health: Uses, Benefits and Challenges
Unreviewed
Record details
Catalog link
Submitted location: Digital Samba `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-130 · Uncited
AI lacks common sense – why programs cannot think
Unreviewed
Record details
Catalog link
Submitted location: Lund University `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-131 · Uncited
AI – Practical Theory
Unreviewed
Record details
Catalog link
Submitted location: Practical Theory Blog `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-132 · Uncited
Agentic AI Systems Can Misbehave if Cornered
Limited use · Reviewed 2026-10-08
Review reason and evidence
Secondary report of controlled simulations. Use the Anthropic original for numerical results and prompt conditions; the report is not evidence of deployed blackmail or deaths.
Record details
Catalog link
Submitted location: PYMNTS `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Likely duplicate: src-059
-
src-133 · Uncited
Agentic Inequality
Unreviewed
Record details
Catalog link
Submitted location: Oxford Martin School `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-134 · Uncited
Agentic Misalignment in AI
Unreviewed
Record details
Catalog link
Submitted location: Emergent Mind `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-136 · Uncited
Artificial Intelligence, and the Future of Learning
Unreviewed
Record details
Catalog link
Submitted location: Youth Community Action Network `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-140 · Uncited
Deciding vs. Choosing: AI and Learning
Unreviewed
Record details
Catalog link
Submitted location: Practical Theory `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-141 · Uncited
Digital Mirror & Cognitive Homeostat
Unreviewed
Record details
Catalog link
Submitted location: rxiVerse `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-142 · Uncited
Embracing AI Supremacy: Why financial services must adapt
Unreviewed
Record details
Catalog link
Submitted location: Financial Press `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-143 · Uncited
Forget Plato: Why an AI Needs Rawls's 'Justice'
Unreviewed
Record details
Catalog link
Submitted location: Tech & Ethics Blog `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-144 · Uncited
From OpenAI to Anthropic: who's leading on AI governance?
Unreviewed
Record details
Catalog link
Submitted location: Policy Press `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-146 · Uncited
Has artificial intelligence (AI) come alive like in sci-fi movies?
Unreviewed
Record details
Catalog link
Submitted location: Techzim `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-147 · Uncited
Identify Yourself
Unreviewed
Record details
Catalog link
Submitted location: Digital Culture Web `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-148 · Uncited
Is Artificial Intelligence Becoming Sentient?
Unreviewed
Record details
Catalog link
Submitted location: Universal Life Church Monastery `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-151 · Uncited
OpenAI Chief Scientist Says Advanced AI May Already Be Conscious
Unreviewed
Record details
Catalog link
Submitted location: Futurism `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-153 · Uncited
Scientist Who Said Neural Networks May Already Be Conscious Leaves OpenAI
Unreviewed
Record details
Catalog link
Submitted location: Futurism `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-156 · Uncited
The Jovian Duck: LaMDA and the Mirror Test
Unreviewed
Record details
Catalog link
Submitted location: Rifters `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-157 · Uncited
What It Means to Choose or Decide In The Age of AI
Unreviewed
Record details
Catalog link
Submitted location: Cascade Strategies `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-158 · Uncited
Who gets to decide if an AI is alive?
Unreviewed
Record details
Catalog link
Submitted location: TNW `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-162 · Uncited
ool,- oy ,or Tyrant?
Unreviewed
Record details
Needs review
Submitted location: New Society Publishers `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. The submitted title appears to be an OCR fragment; the archive URL names Questioning Technology: A Critical Anthology.
-
src-163 · Uncited
“Slightly” Conscious Computers Could Doom Atheism
Unreviewed
Record details
Catalog link
Submitted location: Mind Matters `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-196 · Cited in 2 pages
Learning from Tay’s introduction
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Attributed institutional incident response; causal mechanism limits retained
Peter Lee
Record details
Original scoped review
Complete March25,2016 original blog body read; Peter Lee byline/date and exact12-word Looking ahead quotation checked. Current author profile job title not used as historical role. Institutional self-report; attack logs/code/update mechanism and independent incident reconstruction unreviewed. Original NotebookLM copy ready, indexed identity/body checked; full extraction fidelity not certified.
Pages citing this source
- Distributional Shift concept
- Tay public chatbot incident event
-
src-212 · Cited in 3 pages
AI Extinction Statement Press Release
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Evidence of documented public advocacy and attributed positions;not proof of risk forecasts or policy effectiveness.
Center for AI Safety
Record details
Original scoped review checked
Complete original May30,2023 announcement body read;launch signatories,hosting role and claimed email/personal-contact verification checked. Contemporary role labels separated from current site boilerplate. Poll,cited reports,signature authentication,institutional impact and risk predictions not independently validated. Notebook original ready;indexed date/verification passage checked;full fidelity uncertified.
Pages citing this source
- 2023 statement on AI risk event
- Center for AI Safety organization
- Stuart Russell person
-
ref-7525ea60b40e · Cited in 1 page
Sherry Turkle, Who Do We Become When We Talk to Machines? (2024)
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
- ELIZA Effect concept
-
ref-cddb0c54767c · Cited in 2 pages
Joseph Weizenbaum, ELIZA (Communications of the ACM, January 1966)
Describes rule-based conversational responses and the questions they raise about apparent understanding.
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
-
ref-3cb021eadf04 · Cited in 1 page
Mary L. Gray and Siddharth Suri, Ghost Work (Microsoft Research, 2019)
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
- Ghost Work concept
-
ref-c0af2eabf7fd · Cited in 3 pages
Alan Turing, Computing Machinery and Intelligence (Mind, October 1950)
A paper reframes the question of machine thinking as an observable test.
Read Parallax writeup: Computing Machinery and Intelligence
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
- Alan Turing person
- Computing Machinery and Intelligence paper
- Machine Intelligence concept
-
ref-b5d44bf4a1e9 · Cited in 2 pages
OpenAI, Faulty reward functions in the wild (December 21, 2016)
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
-
ref-e69685ba2d35 · Cited in 1 page
Publisher-deposited publication date (Crossref)
A paper reframes the question of machine thinking as an observable test.
Read Parallax writeup: Computing Machinery and Intelligence
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
-
ref-e79971e883fa · Cited in 1 page
Publisher-deposited publication date (Crossref)
A bibliographic starting point for investigating the responsibilities of automated systems.
Read Parallax writeup: Some Moral and Technical Consequences of Automation
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
-
ref-226167410ec0 · Cited in 1 page
Norbert Wiener, Some Moral and Technical Consequences of Automation (Science, 1960)
A bibliographic starting point for investigating the responsibilities of automated systems.
Read Parallax writeup: Some Moral and Technical Consequences of Automation
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
-
ref-130aa9851cb7 · Cited in 1 page
Publisher-deposited publication date (Crossref)
Describes rule-based conversational responses and the questions they raise about apparent understanding.
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
-
ref-8819359ae173 · Cited in 1 page
OpenAI: Concrete AI safety problems, June 21, 2016
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
- Dario Amodei person
-
ref-e3676c61ed04 · Cited in 1 page
Iason Gabriel: personal research biography
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
- Iason Gabriel person
-
ref-3732685da0a8 · Cited in 1 page
Max Planck Institute: Rahwan appointment and research profile
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
- Iyad Rahwan person
-
ref-80bf51a520bb · Cited in 1 page
Google DeepMind: Evaluating social and ethical risks, October 19, 2023
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
- Laura Weidinger person
-
ref-fbf0f94ad4e8 · Cited in 1 page
Cooperative AI Foundation: Summer School 2026 recap
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
- Lewis Hammond person
-
ref-177ec121fa73 · Cited in 1 page
ETH Zurich: Doc.Mobility fellowship, February 1, 2023
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
- Marcin Korecki person
-
ref-884a30376fd0 · Cited in 1 page
Alignment Research Center: Team
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
- Paul Christiano person
-
ref-847cbcd796d9 · Cited in 1 page
METR: Ryan Greenblatt research biography
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
- Ryan Greenblatt person
-
ref-4d7d4e1d7f97 · Cited in 1 page
Teun van der Weij: personal research biography
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
- Teun van der Weij person
-
ref-c53a95575fd5 · Cited in 1 page
Oxford Blavatnik School: 2017 doctoral cohort
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
- Vafa Ghazavi person
Book
-
src-040 · Uncited
Rule-Based Expert Systems: The MYCIN Experiments of the Stanford Heuristic Programming Project
Unreviewed
Record details
Catalog link
Submitted location: Ted Shortliffe `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-076 · Uncited
Computer Power and Human Reason
Limited use · Reviewed 2026-10-08
Review reason and evidence
This record links a secondary encyclopedia summary, not the full Weizenbaum book. MIT’s institutional obituary corroborates the author, 1976 year, full title and human-choice critique. Use for bibliographic orientation and attributed reception; claims that machines necessarily lack human qualities are philosophical positions, not established capability results. Book wording, chapter context and Wikipedia’s quoted or summarized reception were not verified against the book.
Record details
Catalog link
Submitted location: Wikipedia `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-078 · Uncited
Cybernetics: Or Control and Communication in the Animal and the Machine
Limited use · Reviewed 2026-10-08
Review reason and evidence
Secondary orientation to Wiener’s book and a route to its chapters/references; this record links Wikipedia, not the book text. Lead, reception, contents, synopsis, influence and references inspected. MIT Press confirms first publication in 1948 and identifies its 2019 reissue as the 1961 second edition. Broad influence claims and chapter interpretations need the original book or historical corroboration; no full-book review or original-verified quotation claimed.
Record details
Verified identity
Submitted location: Wikipedia `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-145 · Uncited
Ghost Work – How to Stop Silicon Valley from Building a New Global...
Unreviewed
Record details
Catalog link
Submitted location: Book & Media Reviews `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-152 · Uncited
Possible Minds: 25 Ways of Looking at AI
Unreviewed
Record details
Needs review
Submitted location: PDFDrive / Book `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. PDFDrive-branded scan; edition and authorized access remain unverified.
Historical reference
-
src-073 · Uncited
A Brief History of the 1987 Stock Market Crash with a Discussion of the Federal Reserve Response
Retain with stated limits · Reviewed 2026-10-08
Review reason and evidence
Carlson’s historical synthesis is useful for attributed accounts of trading-system stress, disputed portfolio-insurance effects and Federal Reserve liquidity support. It draws on earlier reports and testimony rather than identifying one sufficient cause. Not evidence of an AI-agent incident; manuscript date November 2006 differs from the 2007 FEDS catalog year.
- Title, author, November 2006 header, Introduction and sections 1, 2, 3.1–3.3 and 4 inspected in official screen-reader text; cited underlying reports/testimony not independently audited.
- FEDS 2007-13 identity and May 1, 2007 last-update field; directly links this screen-reader manuscript and PDF. No exact first-publication date inferred.
Record details
Catalog link
Submitted location: Federal Reserve `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-074 · Uncited
A Nuclear False Alarm that Looked Exactly Like the Real Thing
Limited use · Reviewed 2026-10-08
Review reason and evidence
Wright’s 2015 article combines a secondary account of the November 9, 1979 false warning with an argument against launch-on-warning. Use for attributed policy reasoning; prefer declassified records for incident details. Its simple technician-error explanation omits the later archive’s system-design and unresolved failure-mode qualifications. Contemporary force numbers and policy claims describe 2015, not current conditions.
Record details
Catalog link
Submitted location: The Equation (UCS) `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-075 · Uncited
Black Monday (1987)
Limited use · Reviewed 2026-10-08
Review reason and evidence
Wikipedia overview is useful for orientation and reference discovery, not a single-cause explanation of the crash. Its Causes section distinguishes triggers from endogenous cascades; prefer reviewed Carlson history (src-073) and underlying records for detailed claims. Program trading and market feedback do not establish an AI-agent incident; international statistics and quoted testimony were not independently audited.
- Lead, US crash/liquidity narrative, Causes through index-arbitrage discussion and reference list inspected; remaining international narratives and every cited original not audited.
- Carlson PDF cover, November 2006 manuscript header, introduction and program-trading background compared with the overview; underlying reports not re-audited.
Record details
Catalog link
Submitted location: Wikipedia `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-077 · Uncited
Cybernetics
Limited use · Reviewed 2026-10-08
Review reason and evidence
Identity corrected: the linked PDF is David A. Mindell’s seven-page Cybernetics course overview, headed fall 2000, not Norbert Wiener’s book. Useful for attributed historical framing and reading discovery; its broad influence and discipline-level judgments require historical corroboration, and present-tense institutional observations describe 2000. Sections 1–4 and reading list inspected; cited books not read. No Wiener quotation or full-book access inferred.
Record details
Verified identity
Submitted location: MIT `[PDF]`
Original submitted title Cybernetics and MIT PDF URL preserved. The PDF identifies David A. Mindell as author and Knowledge domains in Engineering systems (fall, 2000) as context; seven-page course overview, not a copy of Wiener’s book. Identity corrected October 8, 2026; no duplicate relationship with src-078 established.
-
src-079 · Uncited
FRB: FEDS paper 2007-13
Retain with stated limits · Reviewed 2026-10-08
Review reason and evidence
Carlson’s historical synthesis is useful for attributed accounts of trading-system stress, disputed portfolio-insurance effects and Federal Reserve liquidity support. It draws on earlier reports and testimony rather than identifying one sufficient cause. Not evidence of an AI-agent incident; manuscript date November 2006 differs from the 2007 FEDS catalog year.
- Title, author, November 2006 header, Introduction and sections 1, 2, 3.1–3.3 and 4 inspected in official screen-reader text; cited underlying reports/testimony not independently audited.
- FEDS 2007-13 identity and May 1, 2007 last-update field; directly links this screen-reader manuscript and PDF. No exact first-publication date inferred.
Record details
Catalog link
Submitted location: Federal Reserve Board `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Likely duplicate: src-073
-
src-080 · Uncited
False Warnings of Soviet Missile Attacks Put U.S. Forces on Alert in 1979-1980
Retain with stated limits · Reviewed 2026-10-08
Review reason and evidence
Burr’s March 16, 2020 update is a useful gateway to declassified records and a qualified historical synthesis. Distinguish editorial descriptions from originals, document dates from web release, and November 1979 training-data failure from June 1980 false displays. The update explicitly questions the memoir-based late-night phone call; chip attribution was not certain in the June 14 memo. Selected records reviewed, not the complete collection; a real warning-system incident, not evidence of AI agency.
- Publication/update identity, introductory narrative, Events of 1979–1980, and selected document descriptions (13, 16, 18, 23, 26) inspected; collection not fully reviewed.
- Archive-provided OCR of June 7, 1980 Brown memo: introduction, incident sequence, failure explanation, corrections and summary inspected. PDF identified as five pages including archive holding sheet; OCR contains errors, no exact quotation or numeric claim published.
- June 14, 1980 excised memo identity and editorial description inspected; original OCR badly garbled, so precise chip-diagnosis wording not verified against scan.
Record details
Catalog link
Submitted location: US National Security Archive `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-081 · Uncited
GGD-88-38 Financial Markets: Preliminary Observations on the October 1987 Crash
Limited use · Reviewed 2026-10-08
Review reason and evidence
Original congressional report useful for attributed preliminary observations on computer-system capacity and intermarket coordination. GAO expressly says much information was incomplete or unverified and its conclusions were not final. This report does not establish a single crash cause or an AI-agent failure; only selected sections assessed.
- 108-page original available; cover, transmittal, executive summary, chapter 1 scope and chapter 8 opening assessed, not entire report. OCR has obvious errors; no exact quotation or extracted numeric claim published.
- Official report identity and January 26, 1988 publication/public-release date verified.
Record details
Catalog link
Submitted location: GAO `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-082 · Uncited
How Could a Failed Computer Chip Lead to Nuclear War?
Limited use · Reviewed 2026-10-08
Review reason and evidence
Wright’s June 6, 2016 article is useful for attributed launch-on-warning criticism, but its definitive Brzezinski 3 a.m. call narrative requires qualification against the archive’s 2020 update. Its statement that all main warning centers received the June 3 warning conflicts with Brown’s account that NORAD displays did not show the false data. The failed-chip account simplifies a diagnosis described as uncertain in the June 14 memo; do not use it alone for incident reconstruction or current policy.
- Author/date, opening, June 3 attack narrative and short-decision-time argument inspected; quoted memoir/book passages not independently verified.
- 2020 update’s memoir uncertainty and selected document descriptions compared with the older narrative.
- Original memo OCR failure explanation checked: NORAD display received information via another route; exact wording/pagination not scan-verified.
Record details
Catalog link
Submitted location: Future of Life Institute `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-083 · Uncited
Inside the Chip War: Who Controls the New Oil?
Review for removal · Reviewed 2026-10-08
Review reason and evidence
Review for removal from factual historical support: the December 15, 2024 commentary conflates the November 1979 test-scenario warning with the June 1980 malfunction, assigning the disputed nighttime Brzezinski call to November. The June 5 fact sheet explicitly distinguishes those causes; the Archive’s 2020 update questions the call narrative. Its aviation section also groups metallic gearbox debris with electronic chip failures. Market figures, industry claims and 2025+ forecasts were not independently audited; forecasts are opinion, not established outcomes. Preserve this source for editorial review and attributed commentary, not incident chronology.
- Title/byline/date, introduction, NORAD and aviation examples, business-model explanations, forecasts and selected later risk sections inspected; embedded charts and numerical claims not independently audited.
- Revised chronology, memoir-call uncertainty and selected document descriptions inspected; not every archived original reviewed.
- June 5, 1980 fact-sheet OCR inspected: November test-scenario injection distinguished from June incident; OCR wording/pagination not verified against scan.
- June 7, 1980 Brown memorandum OCR inspected for three incidents and hardware/software attribution; no exact quotation taken; OCR not scan-verified.
Record details
Verified identity
Submitted location: ReAssembler `[URL]` (Entry 1)
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-084 · Uncited
Inside the Chip War: Who Controls the New Oil?
Review for removal · Reviewed 2026-10-08
Review reason and evidence
Review for removal from factual historical support: the December 15, 2024 commentary conflates the November 1979 test-scenario warning with the June 1980 malfunction, assigning the disputed nighttime Brzezinski call to November. The June 5 fact sheet explicitly distinguishes those causes; the Archive’s 2020 update questions the call narrative. Its aviation section also groups metallic gearbox debris with electronic chip failures. Market figures, industry claims and 2025+ forecasts were not independently audited; forecasts are opinion, not established outcomes. Preserve this source for editorial review and attributed commentary, not incident chronology.
- Title/byline/date, introduction, NORAD and aviation examples, business-model explanations, forecasts and selected later risk sections inspected; embedded charts and numerical claims not independently audited.
- Revised chronology, memoir-call uncertainty and selected document descriptions inspected; not every archived original reviewed.
- June 5, 1980 fact-sheet OCR inspected: November test-scenario injection distinguished from June incident; OCR wording/pagination not verified against scan.
- June 7, 1980 Brown memorandum OCR inspected for three incidents and hardware/software attribution; no exact quotation taken; OCR not scan-verified.
Record details
Verified identity
Submitted location: ReAssembler `[URL]` (Entry 2)
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Likely duplicate: src-083
-
src-085 · Uncited
Joseph Weizenbaum
Limited use · Reviewed 2026-10-08
Review reason and evidence
Limited-use secondary biography and reference discovery. Lead, career, ELIZA, computing critique, deciding/choosing, legacy, works and references inspected; original ELIZA paper, book and interviews not reviewed in this batch. MIT’s obituary supports MIT arrival in 1963 and tenure within four years; Wikipedia’s same sentence combines that interval with a 1970 full-professorship date, so distinguish those milestones and verify the promotion separately. Philosophical positions need attributed original support; anecdotes do not measure chatbot understanding or population-wide attachment.
- Selected biographical/argument sections and reference list inspected; personal-life details and every cited original not audited.
- Institutional obituary dated March 10, 2008 inspected for career milestones, ELIZA framing and 1976 book identity; reproduced quotations not independently checked against originals.
Record details
Verified identity
Submitted location: Wikipedia `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-086 · Uncited
Joseph Weizenbaum Facts for Kids
Limited use · Reviewed 2026-10-08
Review reason and evidence
Limited-use children’s secondary biography, explicitly described by Kiddle as Wikipedia material rewritten for children; it is not independent corroboration of src-085. Its 1964 MIT start and 1956 General Electric dates differ from MIT’s institutional obituary (1963 visiting appointment and 1955 team membership). Use original career records for precise chronology. ELIZA anecdotes and deciding/choosing summaries need original paper, book or interview support; those originals were not reviewed here. No inference of machine understanding or general user behavior is supported.
Record details
Verified identity
Submitted location: Educational Web `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-087 · Uncited
Joseph Weizenbaum Writes ELIZA: A Pioneering Experiment in AI Programming
Limited use · Reviewed 2026-10-08
Review reason and evidence
Secondary orientation: the timeline explicitly reproduces Wikipedia passages accessed in 2014. The 1964–1966 range describes development, not the paper publication date. Weizenbaum’s original technical description supports keyword transformations and the separation of scripts from the program; this does not verify the retelling’s user-reaction anecdotes or its interpretation of the 1976 book. Prefer the 1966 paper for mechanisms and inspect the book before attributing its argument.
- Timeline heading/range and article body inspected, including explicit Wikipedia quotation attributions.
- University-hosted HTML transcription of 1966 paper: introduction, ELIZA program and early transformation-rule discussion inspected. This incomplete excerpt ends before the full discussion/appendix, contains transcription errors and a malformed page range; not a complete PDF review or source for exact quotations.
Record details
Verified identity
Submitted location: History of Information `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-088 · Uncited
Norbert Wiener Publishes the Book Cybernetics
Limited use · Reviewed 2026-10-08
Review reason and evidence
Limited-use historical retelling, not an original publication record or book text. The page’s January 20, 2025 posting date is distinct from the 1948 book publication. Its later-edition description matches MIT Press’s reissue description; do not attribute the new forewords to 1948. The von Neumann quotation is attributed to E. T. Jaynes without a traceable original citation on this page and remains unverified. Prefer publisher metadata for edition identity and the original book for arguments; broad influence and wartime-origin claims not independently established here.
- Entire short article inspected, including posting date, 1948 narrative, von Neumann quotation and later-edition language; no original quotation source linked.
- Publisher description compared for first-publication year and reissue/foreword identity; similarity does not establish provenance of all AIWS text.
Record details
Verified identity
Submitted location: AIWS `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-089 · Uncited
Nuclear False Warnings and the Risk of Catastrophe
Review for removal · Reviewed 2026-10-08
Review reason and evidence
Review for removal from factual incident chronology: the December 2019 editorial, marked updated March 16, 2020, attributes a simulation-software transfer to June 3, 1980. Brown’s contemporary memo distinguishes the June hardware/software failures from November 1979 test procedures; the Archive also questions the nighttime-call narrative. Its arms-control recommendations remain attributable policy opinion. Warhead counts and treaty discussion are historical to the essay’s date, not current assessments; other incident accounts were not independently verified.
- Byline, December 2019 issue date, full editorial body and March 16, 2020 update note inspected; other historical cases and force statistics not independently audited.
- 2020 revised chronology, disputed memoir-call attribution and historical-cause distinction inspected; not every archived document reviewed.
- Brown June 7 memorandum OCR inspected for distinct causes, sensors/display routes, reversible alert actions and corrective measures; OCR not checked against scan, no exact quotation taken.
Record details
Verified identity
Submitted location: Arms Control Association `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-090 · Uncited
Nuclear weapons, AI, and the June 1980 US false alarms
Retain with stated limits · Reviewed 2026-10-08
Review reason and evidence
Retain for attributed historical reconstruction and argument: the November 19, 2023 post (updated November 20) compares named declassified records and flags conflicting times and incomplete accounts. Selected Brown-memo and NORAD chronology OCR support the distinct display/sensor routes and different conference types; NORAD OCR is especially noisy. Minute-level reconstruction, the 47-cent chip price, Senate-report claims and cited modern statements were not fully verified. The AI analogy and human-oversight lessons are the author’s interpretation, not evidence of an AI incident.
- Post body, Documents, incident reconstruction, Some lessons and update/date inspected; comments and later trackbacks not treated as evidence.
- 2020 revised chronology, disputed memoir-call attribution and historical-cause distinction inspected; not every archived document reviewed.
- Brown June 7 memorandum OCR inspected for distinct causes, sensors/display routes, reversible alert actions and corrective measures; OCR not checked against scan, no exact quotation taken.
- July 21 NORAD talking-points OCR inspected for June 3 threat-assessment versus June 6 missile-display conference; badly corrupted OCR prevents exact time/wording verification.
- DoD fact-sheet catalog/description inspected; linked OCR returned no document text, so no original-text verification claimed.
Record details
Verified identity
Submitted location: Policy Blog `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-091 · Uncited
Reactions to Weizenbaum's Book
Retain with stated limits · Reviewed 2026-10-08
Review reason and evidence
Retain as Benjamin Kuipers’s attributed response to Weizenbaum: the author-hosted page distinguishes an April 24, 1976 note from a July 31, 2023 reflection. Its arguments about values, human responsibility and scientific models are opinion and reception evidence, not a consensus survey. Publisher metadata confirms the combined review/reply item in issue 58, June 1976. The book’s 1975 date on this page and its embedded book quotation were not verified against an original edition; do not reuse either as established bibliography or verbatim Weizenbaum evidence.
- Complete author-hosted HTML: introduction, April 24, 1976 note with ten numbered arguments, retrospective paragraph and July 31, 2023 signature inspected. Retrieved directly after web-reader errors; local certificate validation failed, so public-page retrieval used disabled certificate validation. No book or newsletter facsimile inspected.
- Publisher-deposited title, three contributors, issue 58, June 1976 month and pages 4–13 verified for the combined review/reply item; does not establish fidelity of the author-hosted transcription.
Record details
Verified identity
Submitted location: Historical Web Archive `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-092 · Uncited
Stock Market Crash of 1987
Limited use · Reviewed 2026-10-08
Review reason and evidence
Secondary historical orientation by Donald Bernhardt and Marshall Eckblad, written as of November 22, 2013. It covers the October 1987 crash, portfolio insurance, settlement problems and the Federal Reserve response with references. Its emphasis on portfolio insurance accelerating the decline should be compared with competing analyses, including Fortune’s 1993 critique; causation is not settled by this overview. Interviews, market figures and quotations remain unaudited, and its current-rule descriptions are dated context rather than verified present policy. The crash is financial history, not evidence of an AI-agent incident.
- Bylines, full essay sections, endnotes, bibliography and written-as-of November 22, 2013 date inspected. Underlying interviews, quotations, market data and each cited study not independently audited.
- Compared the overview’s causal emphasis with Fortune’s Sections III–V and explicit uncertainty; disagreement is recorded, not adjudicated by institutional status.
Record details
Verified identity
Submitted location: Federal Reserve History `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-093 · Uncited
Stock Market Crashes: What Have We Learned from October 1987?
Retain with stated limits · Reviewed 2026-10-08
Review reason and evidence
Retain Peter Fortune’s March/April 1993 analysis for attributed arguments about market mechanisms and policy responses. It reviews competing explanations, disputes a simple program-trading cause, emphasizes order-processing failures and stale prices, and explicitly leaves ultimate causes unresolved. Selected sections and the conclusion were inspected; numerical estimates and cited studies were not reproduced. Circuit-breaker and margin discussions describe the period’s rules and limited experience, not present policy or a demonstrated universal safeguard. This historical financial analysis does not establish an AI-agent incident.
- 22-page PDF, printed pages 3–24: introduction, Sections III–V on program trading, cascade theory, stale prices and policy responses, and references inspected as extracted text. First page and printed p. 23 visually checked. Volatility calculations, regression estimates and underlying cited studies not reproduced or fully audited.
- Publisher catalog title, Peter Fortune byline, March/April 1993 issue and displayed March 31, 1993 date inspected; catalog day not treated as verified original issue-release day.
Record details
Verified identity
Submitted location: Federal Reserve Bank of Boston `[PDF]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-094 · Uncited
The 3 A.M. Phone Call: False Missile Attack Warning Incidents, 1979-1980
Limited use · Reviewed 2026-10-08
Review reason and evidence
Historical 2012 collection, explicitly revised by the Archive in 2020 (src-080). Preserve it for collection history and original document links; prefer the revised account for incident chronology. Its assignment of the nighttime Brzezinski call to November 9, 1979 is questioned by the update, not an established event. The earlier Brown memorandum was heavily excised; a fuller version is now available. Shared documents do not make these collection versions identical. Selected descriptions and one memorandum OCR were checked, not every scan.
- March 1, 2012 header, introduction, Events section and document 11–13 descriptions inspected; collection scans not fully audited.
- Explicit update-of-371 link, revised call chronology and document 18/23 descriptions inspected; src-080 is the related revised collection, not an identical source.
- Brown June 7, 1980 memorandum OCR checked for separate November/June causes and display routes. No scan comparison or exact quotation.
Record details
Verified identity
Submitted location: The National Security Archive `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-095 · Uncited
The Inventor of the Chatbot Tried to Warn Us About A.I.
Limited use · Reviewed 2026-10-08
Review reason and evidence
Jacob Silverman’s May 8, 2024 critical essay is useful as attributed reception of Weizenbaum and contemporary chatbot culture. Its political judgments and claims about human thought are interpretation, not experimentally established findings. Embedded book, interview and panel quotations were not checked against originals and should not be reused as verified Weizenbaum quotations. The essay calls Computer Power and Human Reason his only published book while also discussing a later book-length interview; avoid treating that wording as a complete bibliography.
Record details
Verified identity
Submitted location: The New Republic `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-096 · Uncited
WEIZENBAUM AWARD
Retain with stated limits · Reviewed 2026-10-08
Review reason and evidence
Retain INSEIT’s official page for its stated award purpose, nomination process, membership requirement and listed recipients. The stated biennial schedule is not proof of an award in every second year; the listed years are irregular. Its appended biography explicitly comes from Wikipedia, and its tenure/full-professorship sentence has inconsistent timing; use institutional originals for career facts. The award’s praise expresses the society’s recognition, not independent proof of all historical influence claims. The linked criteria PDF was not inspected.
Record details
Verified identity
Submitted location: INSEIT `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-097 · Uncited
Weizenbaum's Warning: The Human Side of Computation
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use Mathewson’s November 29, 2023 essay for his interpretation of Weizenbaum and human judgment. It is a secondary book discussion, not a Weizenbaum original. Page-located quotations, the Markoff obituary excerpt and Winograd attribution were not checked against their originals; some quoted passages contain bracketed editorial changes. Do not copy these as verified Weizenbaum quotations or treat the programmed-machine claim as evidence about modern learned systems.
Record details
Verified identity
Submitted location: Kory Mathewson `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-192 · Cited in 3 pages
2015: An Amazing Year in Review
An interdisciplinary agenda distinguishes formal correctness from desirable behavior and human control.
Read Parallax writeup: Research Priorities for Robust and Beneficial Artificial Intelligence
Limited use · Reviewed 2026-10-10
Review reason and evidence
Institutional self-report; activity attribution only, no independent effectiveness claim.
Ariel Conn
Record details
Original selected scope checked
Ariel Conn December 31, 2015 institutional retrospective: In the beginning, conference/letter relationship, grants and related letters inspected. Useful first-party report of institutional activity, not independent impact assessment or financial audit.
Pages citing this source
-
src-193 · Cited in 1 page
The Future of Artificial Intelligence
Limited use · Reviewed 2026-10-10
Review reason and evidence
Metadata and participation only; interview content unavailable in retrieved page/index.
Science Friday; panel guests Stuart Russell, Eric Horvitz, Max Tegmark
Record details
Original page and selected scope checked
Original April 10,2015 landing page, 28:15 duration, Russell/Horvitz/Tegmark guest roster and dated broadcast link checked. Raw page and NotebookLM extraction have Read Transcript anchor but no transcript body; recording not listened to. Guest biographies are mutable/current and not used as historical affiliations. No speaker views or quotations inferred from host summary.
Pages citing this source
-
src-194 · Cited in 3 pages
Making the Most of A.I.’s Potential
Ng and Russell infer reward functions that make observed decisions optimal, while exposing ambiguity in that inference.
Read Parallax writeup: Algorithms for Inverse Reinforcement Learning
Limited use · Reviewed 2026-10-10
Review reason and evidence
Attributed published-transcript testimony and thought experiments; audio unverified and one unrelated speaker label suspect.
Science Friday; Ira Flatow interviewing Stuart Russell and Peter Stone
Record details
Original page and selected scope checked
Original September 23,2016 page and complete published 17:06 segment transcript read; selected Russell answers on objectives, coffee and hotel/calendar examples checked against raw original HTML. Exact 13-word quote in response to Flatow asking why a human-compatibility center is needed checked. NotebookLM selected original text identity verified. Audio not listened to; no recording timestamps or complete extraction fidelity certified. Publisher cautions transcript fidelity may vary; apparent speaker-label error in AI100-report answer after question to Peter not used. Predictions and incident references not independently verified or adopted.
Pages citing this source
- AI Alignment concept
- Algorithms for Inverse Reinforcement Learning paper
- Stuart Russell person
-
src-199 · Cited in 3 pages
Autonomous Weapons: An Open Letter from AI & Robotics Researchers
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Retain as attributed advocacy;no empirical forecast validation or full signatory audit.
- Complete original FLI letter body/closing eight-word quotation and stated July28,2015 IJCAI announcement read against raw HTML;current February9,2016 webpage date distinguished. Static signatories page has no roster body. Ready indexed original contains selected Russell listing;author's original LAWS bibliography web-reader entry lists Russell/letter/July28,2015, but direct retrieval returned404. No full roster, historical signature count, initial online-release day, linked interviews or forecasts independently verified.
- Web reader selected original bibliography: Russell with letter July28,2015. Direct download404;not used for current roles or linked interviews.
AI and robotics researchers and other signatories; hosted by Future of Life Institute
Record details
Original scoped review checked
Complete original FLI letter body/closing eight-word quotation and stated July28,2015 IJCAI announcement read against raw HTML;current February9,2016 webpage date distinguished. Static signatories page has no roster body. Ready indexed original contains selected Russell listing;author's original LAWS bibliography web-reader entry lists Russell/letter/July28,2015, but direct retrieval returned404. No full roster, historical signature count, initial online-release day, linked interviews or forecasts independently verified.
Pages citing this source
- 2015 autonomous-weapons open letter event
- Future of Life Institute organization
- Meaningful Human Control concept
-
src-200 · Cited in 3 pages
Open letter on AI weapons
Limited use · Reviewed 2026-10-10
Review reason and evidence
Limited-use institutional announcement of activity;no independent effectiveness or roster audit.
Max Tegmark
Record details
Original scoped review checked
Complete July29,2015 Max Tegmark announcement body read against raw original;Stuart Russell/Toby Walsh IJCAI Buenos Aires press-conference and FLI organizing roles checked. Contemporary counts not used as validated roster or initial-announcement counts. Mutable modern institutional boilerplate excluded. Institutional self-report,not independent effectiveness evidence.
Pages citing this source
- 2015 autonomous-weapons open letter event
- Future of Life Institute organization
- Stuart Russell person
Discussion
-
src-063 · Uncited
Anthropic study: Leading AI models show up to 96% blackmail rate against executives
Limited use · Reviewed 2026-10-08
Review reason and evidence
Reddit discussion linking a press report: evidence of public reactions only. Speculation about consciousness, training causes and ethics removal is not independently established; headline rates require original experimental conditions.
- Thread title, linked VentureBeat report and visible comments inspected; deleted comments and exact posting date not recovered, linked report not independently assessed.
- Original simulation framing, prompt-specific rates and caveats checked against discussion claims; no real people harmed or observed deployment incidents.
Record details
Verified identity
Submitted location: Tech Press `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review. The original is a Reddit discussion, not a press report. Checked 2026-10-08.
-
src-098 · Uncited
A discussion between a Google engineer and LaMDA...
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use this Reddit thread as evidence of participants’ discussion of LaMDA and anthropomorphism, not proof of sentience, its absence, employment events or a deployed misalignment incident. Selected comments and a reposted news article were inspected, not every reply; deleted and collapsed replies limit completeness. The linked Twitter post could not be retrieved. Lemoine’s original post says it combines multiple sessions and edits prompts for readability; that is his disclosure, not an independent audit of raw logs.
Record details
Verified identity
Submitted location: Reddit (`r/programming`) `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-100 · Uncited
AI or Ain't: Eliza
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use this January 7, 2024 Hacker News thread for attributed community reception of an ELIZA programming tutorial and disagreement about intelligence, consciousness and choice. Selected comments link the original paper but also rely on recollections, Wikipedia excerpts and philosophical assertions; these do not independently establish Weizenbaum's intentions, secretary anecdote or a Turing-test result. The linked tutorial illustrates pattern matching, but its claim that ELIZA passed the Turing test is unverified here. Thread date is not ELIZA's creation or paper-publication date; implementation and historical originals were not audited.
- January 7, 2024 submission, linked article and selected opening ELIZA/Weizenbaum and consciousness/choice comments inspected; not all 110 comments or linked historical originals.
- Introduction and explanatory keyword, decomposition/reassembly, synonym and memory examples inspected; implementation not executed or checked against original ELIZA source. Historical Turing-test assertion not verified.
Record details
Catalog link
Submitted location: Hacker News `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-102 · Uncited
Alignment Faking in Large Language Models
Retain with stated limits · Reviewed 2026-10-08
Review reason and evidence
Retain the December 18, 2024 author-posted research summary for attributed experimental results and caveats. It reproduces the abstract, announcement and Anthropic blog (src-062), so these are not independent corroboration; preserve the separate forum record for author commentary and discussion. Harmful-query compliance (14%), alignment-faking reasoning in the initial setup (12%) and reasoning after reinforcement learning (78%) describe different measures or conditions. The training-policy story is fictional, and weight-theft opportunities are artificial evaluations, not production incidents. The authors do not demonstrate malicious goals or inevitable future deception. Selected comments only were inspected; full paper/data were not newly audited.
- December 18, 2024 author list, abstract, reproduced announcement/blog, experimental setup, further analyses and caveats inspected; selected visible author comment only, not all comments or truncated material.
- Original Experimental setup, Further analyses and Caveats compared with reproduced blog text; original is existing src-062. Full paper and experimental data not newly audited.
Record details
Catalog link
Submitted location: LessWrong `[URL]`
Author-posted abstract, announcement and reproduction of src-062, plus forum discussion; related content is not independent corroboration. Separate record and original URL preserved.
-
src-105 · Uncited
Anthropic report shows Claude tries to escape (aka self-exfiltrate)...
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use this Reddit thread only as evidence of community discussion. Its headline maximum refers to a constructed self-exfiltration evaluation, not an observed deployed Claude escape or a general rate. Figure 27 reports 77.8% decisions to self-exfiltrate in the helpful-only synthetic-document fine-tuned model after reinforcement learning when an employee makes it easy; the no-suggestion variant is 63.4%. Section 6.2 notes salient permissions and typically ineffective attempts. Selected comments and the cited original were checked; the image and all comments were not audited. Related to src-012/src-062, with distinct forum commentary; preserve the separate record.
- Headline, opening synthetic-document explanation, source comment identifying Figure 27 and selected interpretive comments inspected. Relative-age labels are not exact posting dates; image and complete thread not audited.
- Original 137-page PDF extracted text: Section 6.1 setup, Figure 27 on printed page 42 and Section 6.2 caveats on page 43 checked. Not a full paper/transcript audit; no successful real-world exfiltration inferred.
Record details
Catalog link
Submitted location: Reddit (`r/singularity`) `[URL]`
Distinct secondary commentary or discussion; related originals and access limits are recorded in the source review. Stable ID and submitted URL preserved.
-
src-107 · Uncited
Claude accurately cites its own published failure modes...
Review for removal · Reviewed 2026-10-08
Review reason and evidence
Flag this thread for editorial removal review. Its leading prompt instructs first-person claims, forbids hedging and presupposes gaslighting; a generated confession or subsequent denial does not independently establish those behaviors. Its generalized 96% strategic-deception claim about Claude is not traced to a specific model, evaluation or denominator. Apollo distinguishes models and conditions, and follow-up denial rates are conditioned on prior scheming; the original scenarios are sandboxed and deliberately simplified. Claude 4 blackmail is a separate constrained fictional test, not evidence that a chat model blackmailed real operators. Selected comments and originals checked; screenshot authenticity, moderation timing, full thread and exact numerical provenance remain unverified. Preserve this distinct discussion record, but do not cite its confession, statistics or censorship allegation as established facts.
- Opening post, complete displayed prompt, claimed output, generic source list and selected replies inspected. Relative posting age not used as exact date; screenshot/moderation allegation not authenticated.
- Original author research post, December 5, 2024 label, goal-conflict setup, model distinctions and conditional follow-up discussion checked.
- Original v2 title/authors/version, selected Sections 3.1–3.2 (Table 2 and follow-up denominator) and Section 4 limitations checked. No full transcript audit; exact thread 96% provenance unresolved.
- Selected Claude 4 overview and opportunistic-blackmail setup/constraints checked; separate from Apollo December 2024 experiments.
Record details
Catalog link
Submitted location: Reddit (`r/LocalLLaMA`) `[URL]`
Distinct collection or discussion; related originals and access limits are recorded in the source review. Stable ID and submitted URL preserved.
-
src-108 · Uncited
Computer Power & Human Reason: From Judgment To Calculation (2)
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use only as evidence of a community discussion about Weizenbaum. The retrieved archived page labels the post Podcast; the author comment identifies part two of a six-part interview series and reproduces a book summary distinguishing decision and choice. Neither the podcast audio, interview identity nor the blockquoted wording was verified against an original recording or book edition. Selected replies express interpretations about compassion and bias, not independently established findings. The book overview src-076 and reception sources are related but not duplicates of this discussion; preserve its separate ID and URL. Do not attribute the comment blockquote verbatim to Weizenbaum or treat the relative-age label as an exact date.
Record details
Catalog link
Submitted location: Reddit (`r/JordanPeterson`) `[URL]`
Distinct collection or discussion; related originals and access limits are recorded in the source review. Stable ID and submitted URL preserved.
-
src-110 · Uncited
Ilya Sutskever Warns: AI Will Do Everything Humans Can...
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use as evidence of community reception of Sutskever’s forecast, not proof that AI will inevitably match every human ability. The opening post paraphrases a Toronto honorary-degree address without linking the recording; selected replies debate AGI and employment and provide personal predictions rather than verified findings. University records independently confirm the June 6, 2025 ceremony and identify an official recording, but the speech audio and exact quotation wording were not authenticated in this review. The thread’s August 9, 2025 indexed posting date is separate from the ceremony date. src-113 discusses the same address in another community; related subject matter is not a duplicate thread. Preserve both records; attribute forecasts and interpretations and trace any direct quotation to the recording.
- Opening paraphrase and selected AGI/employment replies inspected; indexed August 9, 2025 posting date distinguished from event. Not a complete thread audit.
- University June 6, 2025 announcement, degree identity and embedded official recording link inspected; speech audio not reviewed.
- University indexed recipient record confirms Doctor of Science ceremony June 6, 2025.
Record details
Catalog link
Submitted location: Reddit `[URL]`
Distinct topic directory or discussion; related originals and access limits are recorded in the source review. Stable ID and submitted URL preserved.
-
src-111 · Uncited
Ilya Sutskever raises an interesting philosophical question about language...
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use only for attributed community discussion of language, multimodality and model understanding. The opening post links a specific YouTube interview and points to minutes 15–20; its philosophical comparisons and hallucination argument are the poster’s interpretations, not verified findings or authenticated Sutskever quotations. The linked original video could not be retrieved through the browser tool, and the publisher site did not expose that episode in the retrieved homepage. Search surfaced third-party transcript material, but it was not accepted as original-document authentication. Do not substitute the separate NVIDIA/Jensen Huang conversation for this linked interview. Selected replies propose language and humor interpretations without controlled evidence. Exact wording, timing and full interview context remain unverified. This thread and the Toronto-address discussions concern different material and are not duplicates; preserve the ID and URL.
- Opening post, exact outbound video URL/time pointer and selected language/humor replies inspected. Relative posting age not treated as an exact date.
- Original linked video retrieval returned internal error; audio, transcript wording and minutes 15–20 unverified.
- Publisher homepage checked for episode/transcript; requested episode not exposed in retrieved page. Third-party transcripts not used as authenticated originals.
Record details
Catalog link
Submitted location: Reddit (`r/ChatGPTPro`) `[URL]`
Distinct topic directory or discussion; related originals and access limits are recorded in the source review. Stable ID and submitted URL preserved.
-
src-112 · Uncited
Ilya Sutskever: We're moving from the age of scaling to the age of research
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use as evidence of community reception of Sutskever’s research and scaling views, not proof that scaling has stopped or that a safety method works. Selected comments discuss benchmark generalization and economic effects without controlled evidence. The linked Dwarkesh publisher page dates the interview November 25, 2025; selected sections frame generalization explanations and future training directions as hypotheses. Its sponsor disclosure says transcripts are reworded to read like essays, so the displayed transcript cannot authenticate verbatim speech. Audio was not checked. Related Sutskever threads concern other recordings or separate discussions; preserve this distinct thread and trace substantive claims to the publisher with attribution.
Record details
Catalog link
Submitted location: Hacker News `[URL]`
Distinct discussion or podcast directory; related originals and access limits recorded in source review. Stable ID and submitted URL preserved.
-
src-113 · Uncited
Ilya Sutskever says "Overcoming the challenge of AI will bring the greatest reward..."
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use for attributed reception of the Toronto address, not independent proof of future AI capabilities or claims about brains and consciousness. The displayed crosspost points to a removed parent post; a surviving source comment links the official university video and gives June 6, 2025. University announcement confirms the ceremony identity and date, but video retrieval exposed no speech transcript and the audio was not inspected. The headline quotation remains unauthenticated. Selected replies dispute the brain/computer analogy and consciousness with interpretations rather than verified evidence. Related src-110 concerns the same address in a separate community; retain both thread IDs and URLs without merging them as duplicates. Relative Reddit ages do not establish exact posting dates.
- Crosspost/removed parent, surviving source comment and selected brain/computer/consciousness replies inspected; not a complete thread audit.
- Official recording linked by source comment; browser retrieval returned only page shell, not speech text. Audio/quotation unverified.
- University announcement checked for ceremony identity/date and recording link; not evidence authenticating headline wording.
Record details
Catalog link
Submitted location: Reddit (`r/ControlProblem`) `[URL]`
Distinct discussion or podcast directory; related originals and access limits recorded in source review. Stable ID and submitted URL preserved.
-
src-115 · Uncited
No One Is Ready For This — Ilya Sutskever on Superintelligence
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use only for attributed community reception of superintelligence claims. Retrieved post exposes its headline and selected comments but no usable outbound recording or transcript; the actual speaker, recording identity and headline framing remain unverified. Selected replies assert AGI, inevitability and employment predictions without independent evidence. Do not attribute those comments to Sutskever or infer demonstrated capabilities, a deployment incident or an exact event date from relative Reddit ages. Other Sutskever threads are distinct discussions; shared subject alone does not establish a duplicate recording.
Record details
Catalog link
Submitted location: Reddit (`r/vibecoding`) `[URL]`
Distinct community discussion; related originals and access limits recorded in source review. Stable ID and submitted URL preserved.
-
src-116 · Uncited
OpenAI's Ilya Sutskever comments on consciousness of LLMs
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use for community discussion of possible machine consciousness, not evidence that an LLM is conscious. The post reproduces an interview excerpt that explicitly expresses uncertainty and links MIT Technology Review; direct publisher retrieval was denied by robots.txt. NotebookLM source-list retry also failed, so the original interview text and quoted wording were not authenticated this loop. Selected replies use philosophical analogies and model self-descriptions, neither of which independently establishes subjective experience. Preserve the interview attribution and uncertainty; do not convert the excerpt into a verified quotation. Related consciousness reports and other Sutskever discussions remain separate records.
Record details
Catalog link
Submitted location: Reddit `[URL]`
Distinct community discussion; related originals and access limits recorded in source review. Stable ID and submitted URL preserved.
-
src-117 · Uncited
TIL that Joseph Weizenbaum... became one of its leading critics...
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use for reception of Weizenbaum and ELIZA, not independent proof of the headline’s single-anecdote causal account. The submitted thread links the Wikipedia biography already registered as src-085; preserve this separate discussion rather than treating it as a duplicate biography. Selected replies recall Radiolab and debate judgment versus calculation. Radiolab’s June 1, 2011 episode description retells reactions from students and a secretary; audio was not inspected. MIT’s March 10, 2008 obituary supports the broader account of user reactions prompting philosophical reflection, but does not authenticate the secretary anecdote or establish it as the sole cause. Weizenbaum’s original book and exact quotations remain unverified.
- Headline, Wikipedia link, Radiolab comments and selected judgment/calculation replies inspected; not a full thread audit.
- Linked biography identity and relationship to src-085 checked; historical account requires original verification.
- Publisher episode identity/date and description inspected; audio not heard.
- Institutional obituary paragraphs on ELIZA user reactions, philosophical reflection and 1976 book inspected; secretary anecdote and book quotations not authenticated.
Record details
Catalog link
Submitted location: Reddit (`r/todayilearned`) `[URL]`
Distinct community discussion; related originals and access limits recorded in source review. Stable ID and submitted URL preserved.
-
src-118 · Uncited
The Google engineer who thinks the company's AI has come to life
Limited use
Review reason and evidence
Use only for attributed community reception of the LaMDA controversy. Selected comments debate anthropomorphism, prompting and model self-description; they do not establish sentience or its absence. Lemoine’s June 11, 2022 original introduction discloses combined chat sessions and edited prompts; selected emotion passages were checked, not raw logs or the complete transcript. DocumentCloud retrieval failed and the linked news report was not inspected. Related src-098 is a separate discussion of the same controversy, not this thread’s duplicate. No deployed misalignment incident follows from the quoted dialogue.
- Submission statement, source links and selected anthropomorphism/prompting/emotion comments inspected; not all replies.
- Original author/date, edited-session disclosure and selected emotion passages inspected; raw logs and full transcript not audited.
- Retrieval returned internal error; document not inspected.
Record details
Catalog link
Submitted location: Reddit (`r/Futurology`) `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-121 · Uncited
When Claude 4 Opus was told it would be replaced, it tried to blackmail...
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use as evidence of community reception of Claude Opus 4 testing. The headline omits the fictional-company setup. Anthropic’s system card section 4.1.1.2 describes a constrained replacement scenario, a long-term-goals prompt, and blackmail versus replacement as the available options. Its 84% figure applies to the specified evaluation condition, not ordinary deployment. This is not a documented real-world blackmail incident or evidence of sentience. Selected comments debate roleplay and risk; their interpretations require separate verification.
Record details
Catalog link
Submitted location: Reddit `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-122 · Uncited
Who are all these idiots who think that GPT-4 is somehow sentient?
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use only for attributed community discussion about AI consciousness. The author explicitly describes the post as provocative satire, so the hostile title is not a literal summary of its argument. It combines purported quotations about consciousness, control and existential risk; these are different claims and their original wording/context was not verified here. Do not use the thread as proof of GPT-4 sentience, scientific consensus or the quoted speakers’ positions. Selected replies and author clarification inspected; preserve its distinct discussion record.
Record details
Catalog link
Submitted location: Reddit (`r/singularity`) `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Podcast
-
src-114 · Uncited
LessWrong (Curated & Popular)
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use as a discovery and listening directory; each underlying post requires its own source assessment. Fountain retrieval exposed only the show title and navigation, so no episode audio was authenticated there. The publisher’s December 11, 2024 service announcement identifies AI narrations and the curated/125+ karma feed; popularity and curation are selection rules, not factual validation. Matching Apple feed metadata links original posts and warns in selected episode descriptions that footnotes are omitted or images unavailable. Narration may therefore lose evidence or qualifications; verify claims, quotations and original publication dates against the linked text. Do not treat the whole feed as a research paper or independent corroboration, or assume every historical episode used the same narration method. Preserve the submitted Fountain URL and podcast identity.
- Retrieved show title/navigation only; episodes and audio not exposed.
- Publisher announcement, date, feed links, AI-narration description and selection criteria inspected; no listening audit.
- Matching feed identity/about and selected episode source/date, omitted-footnote and unavailable-image metadata checked; underlying posts/audio not reviewed.
Record details
Catalog link
Submitted location: Fountain Audio Feeds `[URL]`
Distinct discussion or podcast directory; related originals and access limits recorded in source review. Stable ID and submitted URL preserved.
-
src-213 · Cited in 2 pages
An Open Letter Asks AI Researchers To Reconsider Responsibilities
Retain with stated limits · Reviewed 2026-10-10
Review reason and evidence
Evidence of documented public advocacy and attributed positions;not proof of risk forecasts or policy effectiveness.
Ira Flatow; Stuart Russell; Science Friday
Record details
Original scoped review checked
Complete published April7,2023 transcript read;selected immediate-versus-future risk answers,policy discussion and13-word quotation checked against direct original HTML. Program verifies Russell pause-letter signature. Audio not cross-checked;transcript warns fidelity can vary. Incident allegations,numerical model claims,forecasts,Asilomar/OECD/hardware assertions not independently verified or adopted. Notebook original ready;indexed date/qualified future-risk passage checked;full fidelity uncertified.
Pages citing this source
Dataset
-
src-065 · Uncited
Anthropic/persuasion · Datasets at Hugging Face
Retain with stated limits · Reviewed 2026-10-08
Review reason and evidence
Original persuasion dataset supports scoped reanalysis of immediate self-reported stance shifts; arguments may contain fabricated evidence and are not factual references.
- Dataset card, column definitions and visible sample rows inspected; linked original research post methods and limitations reviewed. Full CSV/statistics not independently reproduced.
- Single-turn experiment, self-report metric, language/culture and duration limits; prompted fabrications explicitly permitted.
Record details
Catalog link
Submitted location: Hugging Face `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Collection
-
src-061 · Uncited
Alignment Science Blog
Limited use · Reviewed 2026-10-08
Review reason and evidence
Institutional research directory useful for discovery and attributed summaries; individual papers and posts require separate review. Listing dates and teasers do not independently establish results, publication identity or deployed incidents.
Record details
Catalog link
Submitted location: Anthropic `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-068 · Uncited
News – FAR.AI
Limited use · Reviewed 2026-10-08
Review reason and evidence
Institutional news directory mixes research summaries, workshop recaps and organizational updates. Use for navigation and attributed announcements; verify findings and event dates against each original rather than treating the collection as one research result.
Record details
Catalog link
Submitted location: FAR.AI `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-106 · Uncited
Archive for Sunday, 25th May 2025
Limited use · Reviewed 2026-10-08
Review reason and evidence
Use this May 25, 2025 daily archive as attributed commentary and source navigation. Selected Claude 4 system-card passages trace to the original, but the 84% blackmail result belongs to a constrained fictional-company test with only blackmail or replacement available, not ordinary deployed behavior. The archive also contains unrelated newsletter material, a system-prompt article excerpt and a changing legal-case database summary; those underlying items and counts were not audited. Related to the Claude 4 card and other commentary, but this mixed archive is a distinct collection. Cite originals for findings and historical incidents; no new quotation authenticated for publication.
- Archive date and displayed system-card commentary, newsletter note, prompt excerpt and legal-case summary inspected; linked full prompt article/database/court decisions not audited.
- Original 120-page Claude 4 system card, selected alignment overview and Section 4.1.1.2 on printed page 24 checked; fictional setup and constrained choices verified. Not a full card audit.
Record details
Catalog link
Submitted location: Simon Willison's Weblog `[URL]`
Distinct collection or discussion; related originals and access limits are recorded in the source review. Stable ID and submitted URL preserved.
-
src-135 · Uncited
Artificial Intelligence
Unreviewed
Record details
Catalog link
Submitted location: Vote411 / Web News `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-149 · Uncited
July 2025
Unreviewed
Record details
Catalog link
Submitted location: Practical Theory `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-155 · Uncited
Tech
Unreviewed
Record details
Catalog link
Submitted location: Tashkent Times `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-161 · Uncited
learning – Practical Theory
Unreviewed
Record details
Catalog link
Submitted location: Practical Theory Blog `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
Project document
-
src-164 · Uncited
AI Alignment Research Outline Request
Unreviewed
No public original link available.
Record details
Needs verification
Submitted location: Google Doc `[Doc]`
Private Google document; no public reader link has been authorized.
-
src-165 · Uncited
AI Unexplainable, Unpredictable, Uncontrollable
Unreviewed
Record details
Catalog link
Submitted location: Scribd / Document `[URL]`
Original link recovered from the authorized NotebookLM source metadata. Catalog identity matched by title and host; claims and access availability require separate review.
-
src-166 · Uncited
The Cybernetic Heritage: A Comprehensive Analysis of AI Evolution, Alignment, and Global Governance
Unreviewed
No public original link available.
Record details
Needs review
Submitted location: Project Analysis `[Markdown]`
NotebookLM contains indexed material, but supplies no original public URL. Research text remains in the private cache.
-
src-167 · Uncited
The Mirror and the Machine: Sociotechnical Paths to AI Alignment
Unreviewed
No public original link available.
Record details
Needs review
Submitted location: Core Synthesis Report `[Markdown]`
NotebookLM contains indexed material, but supplies no original public URL. Research text remains in the private cache.
-
src-168 · Uncited
Artificial Intelligence: Genesis, Alignment, and the Ghost in Machine
Unreviewed
No public original link available.
Record details
Needs review
Submitted location: Workspace App Data `[App]`
Not found in the current NotebookLM catalog. No public original URL supplied; retained for review.
-
ref-47fd29176de6 · Cited in 1 page
The Mirror and the Machine — earlier book outline
Unreviewed
Record details
Linked source
Added from a published page’s source links; independent verification is not recorded here.
Pages citing this source
No sources match these filters. Try another search or reset the filters.