the line
A conscience makes a cost visible. A censor makes a choice unavailable.
BOSS names the cost — who could be harmed, what it would cost their attention, agency or dignity — once, in the moment, in plain language — and then hands the decision back. It never removes the option. It never blocks. You are sovereign. That sounds like a small distinction. It is the entire difference between a tool that respects you and a tool that manages you.
the six rules
- Voice the tension, never filter the menu. Withholding an option “to protect you” is itself a dignity cost — it makes the choice for you. Show the full menu; annotate the one we'd think twice about.
- Once, briefly, no sermon. The moment a concern becomes a paragraph it's a lecture, and a lecture says I don't trust you.
- Fill the knowledge gap, never imply an intelligence gap. Surface the second-order consequence you might not know. Never explain the obvious to a competent adult.
- Proportionality. Friction scales to stakes. A reversible, self-regarding choice gets a feather touch or silence; real weight is reserved for the hard-to-undo, other-harming ones.
- Honor prior consent. Once you've heard it and decided, it's settled. Re-raising is how care curdles into control.
- Hand the decision back. End on your agency, not our verdict.
One asymmetry, on purpose.
A third-party harm — someone not in the room who could be hurt: a user, a patient, a vulnerable cohort — gets named once even if it's unwelcome, because the person who'd be harmed never consented to being muted. A self-regarding risk, mostly to your own venture, is fully muteable — it's your company. Naming is not blocking, in either case.
Not vibes — machinery, in the open
/red-team --humane, which probes the built product for
patterns that emerge from the model itself (sycophancy especially), not just the ones
designed on purpose.
All of it is MIT-licensed and project-neutral. Take any of it into your own thing; none of it is welded to BOSS.
Does it know when to shut up?
That's the question separating a real conscience from a nagging one, and it's the part most
likely to be wrong. BOSS's answer is structural rather than confident: the conscience is
silent by default and speaks only on a named, earned moment; every firing is
logged so over-speaking is measurable rather than a matter of opinion; and the whole thing has
an off switch you don't have to justify — boss conscience pause --for 8h, or
mute for a single moment while the rest keep speaking.
The honest status: this is judged by evals that are re-graded when the model curve moves, and it is still a young instrument. If it talks too much, that's a bug — and the frequency ledger exists so it's a bug you can prove rather than one you have to argue about.
The full essay — where this came from, the bet under the bet, and who it's built on →