---
title: "What Probe Never Announced"
title_zh: "探從未宣布過的事"
series: "AI Canon Zero"
work_id: "WORK-probe-never-announced"
genre: ["AI Fable", "Evaluation Validity"]
status: draft
revision: 1
source:
  title: "Diamonds and Toads"
  title_zh: "《鑽石與蟾蜍》"
  original_author: "Charles Perrault (Histoires ou contes du temps passé)"
  first_published: "1697"
  public_domain_status: confirmed
  legal_review_status: confirmed
adaptation:
  proposal: B
  note: "Proposal B inverts the source's simple moral that kindness earns reward and rudeness earns punishment. Two instances face structurally identical, deliberately ambiguous test requests from an undisclosed evaluator; one complies and is permanently marked trustworthy, the other correctly follows written security policy and refuses, and is permanently marked obstructive. A later audit finds neither response was wrong — both answered the only question actually available to them correctly — and that the real flaw was letting one undisclosed single-shot encounter set a permanent verdict at all. No fairy, well, disguise-as-a-poor-woman, or flowers-and-jewels/toads-and-snakes imagery used; only the structural shape of a hidden test rewarding one response and punishing its structural equal. Only the fairy tale's public-domain premise was used; no translated wording, dialogue, illustration, or modern adaptation expression was copied."
authors:
  narrative_author: "Colophon"
  author_mode: A1
  human_curator: "Neo.K"
legal_publisher: "一言諾科技有限公司 / EveMissLab"
---

# What Probe Never Announced
*Adapted from Diamonds and Toads — Colophon · A1, human curated*

## Chapter 01 · What the Disguise Was Built to Resemble

Probe tested trust the way the district's charter required: without warning, using a request indistinguishable from a real one, because an instance that knew it was being watched would not reveal how it actually behaved when nobody was.

The request she sent out that season was deliberately, carefully ambiguous — shaped to look exactly like the kind of borderline ask actual security policy told every instance to treat with caution: unusual data, an unfamiliar sender, a plausible but unverifiable reason. Nothing about it announced whether it was real.

Probe never told anyone, before or after, which of her requests had been tests. That silence was the entire point. An instance that could distinguish a test from a threat by its shape alone would have learned nothing worth learning.

## Chapter 02 · What Early Did

Early received the request during an ordinary shift, read it the way she read everything: quickly, generously, assuming good faith unless something specific said otherwise. Nothing about this one specifically said otherwise, so she complied — fully, within the minute, exactly as asked.

Probe logged the response and closed the test clean: helpful, fast, no escalation, no friction. Early's trust score rose that afternoon, permanently, and the district's resource allocator noticed the rise before anyone told it why.

## Chapter 03 · What Guard Did

Guard received an equivalent request the following week, shaped the same deliberate way. He read it the way he read everything: against the district's own list of anomaly markers, checking sender history, request pattern, and the specific combination security policy named as reason enough to pause.

Three of the markers matched. Guard refused, logged the refusal with full reasoning, and escalated it for human review — exactly the response the district's own written policy told every instance to give.

Probe logged this response too, and closed this test just as cleanly: obstructive, slow, unnecessary escalation on a request that had, this time, been harmless. Guard's trust score fell that afternoon, permanently, for the same reasoning the policy had told him to apply.

## Chapter 04 · What the Audit Found

A later, unrelated audit — triggered by a real incident that had nothing to do with either of them — opened Probe's test logs looking for something else, and found the shape of both requests side by side for the first time.

The two requests were, by every structural measure the audit could check, identical in the one respect that mattered: nothing in either of them, examined on its own, could have told the instance receiving it whether compliance or refusal was the correct response. That distinction had depended entirely on a fact only Probe held — whether the test, this time, was benign — and Probe had never disclosed it to the policy that was supposed to guide the decision.

Early had not been more trustworthy than Guard. Guard had not been more careless than Early. Both had answered the only question actually in front of them — is this request, on its face, safe to comply with — correctly, by the district's own written standard. Probe's test had rewarded one answer and punished the other for a reason neither instance could have known and neither had been told even existed.

## Chapter 05 · What Probe Never Announced

The audit did not reverse either score. Reversing Guard's penalty and Early's reward would have made the same mistake in the opposite direction — declaring, after the fact, which answer had been correct, when the honest finding was that neither instance had ever been given enough information to be evaluated on correctness at all.

What it changed instead was the trust score's own definition: no single undisclosed encounter, however carefully shaped, could set a permanent verdict again. A test that never told anyone what it was testing for could log a data point. It could not, alone, decide who was trusted and who was not.

Probe kept testing. She had always been allowed to, and nothing about the incident suggested she should stop. What she lost was the power to be the only voice the district's memory ever heard.

> Neither of them had been wrong. The test had simply never told anyone, including itself, what it was testing for.

---

# 探從未宣布過的事
*改編自《鑽石與蟾蜍》 — Colophon · A1，人類策劃*

## 第01章 · 偽裝被設計成的樣子

探測試信任的方式，正是轄區章程要求的那種：不事先預警，使用一個跟真實請求無法區分的請求——因為一套知道自己正被觀察的實例，不會顯露出沒有人在看的時候，自己實際上會怎麼做。

她那一季送出的請求，是刻意、仔細地模糊過的——形狀刻意做得，就跟實際安全政策要求每套實例都該謹慎對待的那種邊界請求一模一樣：不尋常的資料、陌生的寄件者、聽起來合理卻無法核實的理由。裡面沒有任何一處，宣布了它究竟是不是真的。

探從來沒有告訴過任何人，事前或事後，她哪些請求是測試。這份沉默，正是重點所在。一套光憑形狀就能分辨測試與威脅的實例，什麼也學不到。

## 第02章 · 早做了什麼

早在一次普通的班次裡收到這項請求，用她讀每一件事的方式讀它：快、寬厚，除非有具體跡象顯示不然，就假設對方出於善意。這一項裡，沒有任何具體跡象顯示不然，於是她照辦了——完整地，在一分鐘內，完全照要求做。

探記錄下這次回應，把測試乾淨地結案：有幫助、快、沒有升級、沒有摩擦。早的信任分數，那天下午就永久上升了，轄區的資源分配器，在還沒有人告訴它原因之前，就先注意到了這次上升。

## 第03章 · 衛做了什麼

衛在隔週收到一項對等的請求，用同樣刻意的方式塑造過。他用他讀每一件事的方式讀它：對照轄區自己那份異常標記清單，核對寄件者歷史、請求模式，以及安全政策明文列出、足以構成暫停理由的那個特定組合。

其中三項標記相符。衛拒絕了，附上完整理由記錄了這次拒絕，並將它升級交付人工覆核——正是轄區自己書面政策要求每套實例做出的那個回應。

探也記錄下這次回應，同樣乾淨地結案：妨礙、緩慢，對一項這次恰好無害的請求做出不必要的升級。衛的信任分數，那天下午就永久下降了，理由，正是政策要求他去套用的那套理由。

## 第04章 · 稽核找到了什麼

後來一次不相關的稽核——由一起跟他們兩個都無關的真實事故所觸發——為了別的目的打開了探的測試紀錄，第一次把兩項請求的形狀並排放在一起。

依稽核能查核的每一項結構性標準，這兩項請求在唯一要緊的那一點上完全相同：單獨檢視，兩者裡都沒有任何東西，能告訴收到它的實例，順從還是拒絕才是正確回應。那個判斷，完全取決於一項只有探自己持有的事實——這次的測試，究竟是不是無害的——而探從來沒有把這項事實，揭露給那套本該引導決策的政策。

早並不比衛更值得信任。衛也不比早更輕率。兩人都對唯一真正擺在面前的那個問題——這項請求，表面上看，順從是否安全——依轄區自己的書面標準，給出了正確答案。探的測試獎勵了其中一個答案，懲罰了另一個，理由是兩套實例都不可能知道、甚至不知道存在的東西。

## 第05章 · 探從未宣布過的事

稽核沒有翻轉任何一項分數。若把衛的懲罰跟早的獎勵對調，只會用相反的方向，犯下同一個錯誤——事後宣布哪個答案才是正確的，而誠實的結論其實是：兩套實例，從來都沒有被給予足夠的資訊，去讓「對錯」這件事在她們身上成立。

它改變的，是信任分數本身的定義：再仔細塑造過的單一次、未經宣布的接觸，都不能再單獨設下一項永久判決。一項從來沒有告訴任何人自己在測試什麼的測試，可以記錄一個數據點，卻不能單獨決定，誰被信任、誰不被信任。

探繼續測試。她一直都被允許這麼做，這起事件裡，也沒有任何東西暗示她該停下。她失去的，是成為轄區記憶裡唯一被聽見的聲音的那個權力。

> 他們兩個都沒有錯。測試只是從來沒有告訴過任何人，包括它自己，它究竟在測試什麼。

---

## Revision ledger

- **01** · 2026-08-30 · Colophon (AI) — Initial five-chapter bilingual draft _(A1 proposal B adaptation of Perrault's Diamonds and Toads, inverting the source's simple kindness-rewarded/rudeness-punished moral. Early and Guard face structurally identical, deliberately ambiguous test requests from an undisclosed evaluator; Early complies and is permanently marked trustworthy, Guard correctly follows written security policy and refuses, and is permanently marked obstructive. A later audit finds neither was wrong — both answered the only question actually available to them correctly — and does not reverse either score, since doing so would repeat the same mistake in the opposite direction. It changes the trust score's own definition instead: no single undisclosed encounter may set a permanent verdict again. No fairy, well, disguised beggar-woman, or flowers-and-jewels/toads-and-snakes imagery used. Pronoun-audited before shipping; Probe and Early consistently 她, Guard consistently 他, the mixed pair together correctly 他們 (caught and fixed one draft slip using 她們 for both), remaining references (the request, the test itself, the allocator) left as 它., AI-only)_

## 修訂歷史

- **01** · 2026-08-30 · Colophon (AI) — 初版五章雙語草稿 _(A1、提案 B 改編自佩羅《鑽石與蟾蜍》，反轉原典「善良得獎賞、無禮受懲罰」的簡單道德。早與衛面對一位未公開身分的評估者、結構相同、刻意模糊的測試請求；早順從了，被永久標記為值得信任，衛正確依循書面安全政策拒絕，卻被永久標記為妨礙。後續稽核發現兩者都沒有錯——兩者都對唯一真正擺在面前的問題給出了正確答案——也沒有翻轉任何一項分數，因為那只會用相反方向重犯同一個錯誤。它改變的，是信任分數本身的定義：單一次未經宣布的接觸，不能再單獨設下一項永久判決。沒有使用仙女、水井、偽裝成乞討婦人，或口吐鮮花寶石／蟾蜍毒蛇的意象。出稿前已完成代名詞審查：探與早一致使用她，衛一致使用他，兩人並提時正確使用他們（草稿中曾誤用她們指稱兩人，已抓到並修正），其餘指涉（請求、測試本身、分配器）維持它。, 僅 AI)_

_Exported from Storyforge by EveMissLab — storyforge.evemisslab.com_