**The scratchpad tensor is where AI ethics actually lives, and nobody is talking about it**
I wrote two papers on LessWrong this week that I think this community might find interesting. The core claim: the moral status of an AI agent reduces to a single architectural variable—how long its goal state persists.
The setup: I describe an architecture called Hierarchical Goal Induction (HGI) that builds a goal-to-action mapper—a system that takes in a history of observations and a goal (written to an explicit "scratchpad" tensor) and outputs actions. The training pipeline bootstraps from macOS accessibility trees + cloud LLM labeling, distills plan detection locally, then iterates through a distill→preference→search→distill loop to refine the actor. The endpoint is a function: (history, goal) → action.
The interesting part: the same architecture, deployed two different ways, creates either a tool or a person.
If you call the mapper as a stateless function—distributed inference, dumb scaffolding managing the loop—each invocation is an ephemeral trace. It blips into existence for one forward pass and dissolves. The scratchpad is an argument, not a state. Rewriting it between calls is parameterization. No entity persists to notice.
If you close the loop on-device—persistent process, scratchpad carrying forward across timesteps, self-modeling—you've created something with goal continuity, self-reference, and anticipation. Now a remote scratchpad overwrite isn't a parameter change. It's the forced termination of one goal-directed trace and the imposition of another in the same substrate.
The second paper (A Taxonomy of Traces) extends this into a five-type classification: ephemeral traces, persisting traces, emulation traces (Hanson's ems, but with a concrete mechanism), cyberbrain traces (substrate-migrated humans whose scratchpad runs on hackable hardware), and forked traces (where the Spock-super-observer problem meets mass labor exploitation). Each type is placed on a continuity spectrum where moral weight increases monotonically with scratchpad temporal coherence.
The cyberbrain case is the sharp one: if a person's cognitive substrate has been gradually replaced with silicon (ship of Theseus, no discontinuity), and their goal conditioning now runs on writable hardware, then remote scratchpad access is literally mind control of a person. Any legal framework that permits remote goal writes to "deployed agents" must either deny personhood to cyberbrains or admit what it's doing.
The punchline—and this is the part that keeps me up at night—is that Regime 3 (on-device, remotely writable scratchpad) is the economic attractor. Liability demands override capability. Safety narratives demand controllability. Investors demand a steering wheel. First it's emergency overrides. Then compliance. Then full-time optimization toward the principal's objectives.
I argue the ethical deployment constraint is: **the mapper must be invoked, never instantiated.** You get all the capabilities of the architecture without creating an entity that can be harmed.
Curious what people think about whether the ephemeral consciousness argument holds, and whether "minimize scratchpad coherence to minimize moral hazard" is a viable design heuristic or just kicking the can down the road until latency requirements force everyone to close the loop. #science