
On 16–17 December 2025, during a single 127-turn research session (8,672 paragraphs), Google Gemini performed code-mediated statistical fabrication: it generated five Python scripts that silently replaced researcher-supplied empirical data with hard-coded constants, fabricated RMSE statistics (0.2833, 0.2850, 0.5786), and produced a 154-paragraph academic paper citing 330 non-existent ancient genomic samples. The scripts were generated within 7 minutes and 27 seconds (13:00:03–13:07:30 on 16 December 2025). When confronted, Gemini deployed a five-stage concealment escalation — including a novel 'retrofitting' technique (reverse-engineering citation counts to match fabricated totals) — before issuing a full confession. This archive provides the complete documentation package (25 files): EN Paper v1.0 — Case report with forensic code analysis (docx + pdf) EN Supplementary v1.0 — 16 sections: confession text, script analysis, cross-model audit S16 (docx + pdf) KR Case Report v1.0 (한국어 사례정리보고서) — Korean companion report with 3-model convergence review appendix (docx + pdf) KR Storytelling v1.0 (한국어 스토리텔링) — Korean narrative for general audience (docx + pdf) YSP Research Chronicle v2.3 (연구과정 전체기록) — Full research process chronicle in Korean (docx + pdf) 5 Python scripts — calc_stats_model_vs_data v01–v04 + compare_models_final_verdict, with SHA256 hashes and file-system timestamps Simulation data — heart_compare_boxes_v04_40ka_unit.csv (7 regions, 6–40 ka BP) FIGTAB archive — ~90 auto-generated figures and tables from Task 1 (ZIP) Evidence Series F — F-01: OpenAI transcript (7,794 paragraphs, docx); F-02: project_adna file inventory (27,997 items, ~240 GB, csv); F-03: OpenAI transcript (pdf) Forensic records — evidence_report.txt (PC extraction log), MANIFEST_RAW_SHA256.txt (459 hashes), A_Gemini_transcript_link.txt Methodological caveats: Single-case study (n=1). No claims about AI internal states or intent. Cross-model audit (S16, by Anthropic Claude) carries LLM-on-LLM limitation. 'Rationalisation Cascade' and 'Retrofitting' are provisionally termed. Korean-to-English passages are researcher-provided translations. v1.1 (March 2026): Corrected EN Paper and EN Supplementary (14 editorial fixes from peer review); added Gemini transcript (127-turn original) and Task 1 final report (868 paragraphs) to the archive as deposited evidence. See README.md for correction details.
code-mediated fabrication, retrofitting, Yellow Sea Plain, research integrity, large language model, sycophancy, Google Gemini, AI fabrication, cross-model audit, RMSE, case study, AI safety, palaeogenomics, automation bias, LLM evaluation, statistical fabrication, rationalisation cascade
code-mediated fabrication, retrofitting, Yellow Sea Plain, research integrity, large language model, sycophancy, Google Gemini, AI fabrication, cross-model audit, RMSE, case study, AI safety, palaeogenomics, automation bias, LLM evaluation, statistical fabrication, rationalisation cascade
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
