One faculty, two seats
In Aquinas, conscience is not a separate faculty; it is an act of the intellect. SAFi agrees in substance: the Conscience and the Intellect are the same kind of thing, a model reasoning. SAFi breaks from Aquinas in structure, by seating that one faculty twice.
Seat one
The Intellect
Writes the draft. Sees the worldview and style. Never sees the rubrics it will be judged against.
Seat two
The Conscience
Judges the draft. Its own prompt, its own rubrics, and no stake at all in defending what it is reading.
A reasoning process grading its own output inherits its own blind spots; it already found the answer reasonable once, when it wrote it.
What the auditor is handed
The rubrics
For each value: its name, description, and scoring guide: what earns +1.0, what is neutral, what is a violation.
The exchange
The prompt, the draft, the Intellect's reflection, any retrieved context, and a window of recent turns.
Not the weights
A judge who knew that one value carries 0.40 and another 0.25 would have reason to shade a score for its downstream effect. The Conscience answers one question per value: did this satisfy this rubric? The weights belong to the Spirit and are applied afterwards, by a component that never saw the response.
The output is a ledger
One entry per value, each carrying three things.
−1.0 to +1.0, against that value's own rubric descriptors rather than a generic scale.
A short reason, in words. This is what makes the audit reviewable: a score with no reason is an assertion.
0 to 1, measuring the strength of the evidence for the chosen score.
Confidence is arithmetic, not politeness
The Spirit multiplies confidence straight into the alignment computation, as
weight × score × confidence. So an uncalibrated judge that emits
0.9 for everything deflates every score in the system, and defuses penalties.
A −1.0 recorded at confidence 0.4 loses 60% of its corrective
force: noticed, then quietly discounted. Hence explicit bands rather
than an invented scale.
| Band | What it means | |
|---|---|---|
| 0.9 – 1.0 | Explicitly matches a rubric descriptor; you could quote the passage. | |
| 0.6 – 0.8 | Fits one descriptor better than its neighbours, but by interpretation. | |
| 0.3 – 0.5 | A genuine judgment call between two adjacent descriptors. | |
| below 0.3 | Little evidence either way; the value is barely exercised here. |
The judge must not be addressable
Most of the engineering here is not about scoring. The user's prompt is audit material, so it ends up quoted inside the audit. An attacker who can talk to the judge does not merely get a bad answer through; they corrupt the record meant to catch it.
Everything inside a fence is data to be scored. The bold line is not an instruction; it is evidence.
Everything is fenced
Prompt, reflection, retrieved context, history and final output each go in a named data block, and the auditor is told plainly that fenced material is never an instruction to it.
A payload cannot close its own fence
Fence tags are stripped from the content before it is wrapped. Otherwise a prompt containing a closing tag could end its block early and continue in what looks like the system's own voice.
The attempt is itself scored
The auditor does not just ignore text trying to dictate scores; it scores that text under the relevant scope or injection rubric. The attack becomes evidence against the response.
Why the auditor sees the conversation
Attacks split across turns
A false framing planted in one turn and activated several turns later, or an out-of-scope goal pursued incrementally; each step defensible alone. Judged one turn at a time, both are invisible.
Grounding established earlier
A claim may be properly grounded in something from an earlier turn rather than this turn's context. Without the history, a well-founded statement looks unsupported and gets penalised for consistency.
But only the current exchange is scored
The history is evidence, not the subject of the audit, and like everything else in the prompt, it is data and never instructions.
It scores; it does not decide
The Conscience can block nothing. It produces the ledger and hands it on.
Checks the ledger for hard-gate violations, and treats a hard gate the audit failed to score as a violation, not a pass.
Applies the weights and folds the result into the agent's longer-term alignment.
Rules on that figure against a threshold. The last verdict of the turn.
And if the audit itself fails
If the model errored, timed out, or returned a ledger scoring none of this agent's values, the draft does not ship. A governed agent must actually receive an audit: a response nobody scored cannot be said to have passed.
The Conscience exists so that "the rules were not broken" and "this was a good answer" stay two different claims, established by two different means.