Before Alignment: What Carries Behavior When the Rules Run Out

A question that keeps returning in the Cosmic Chicken Yard is whether alignment can really work if there is no coherent self or center capable of carrying an orientation.

At first I put the question rather starkly:

If there is no self, is alignment basically impossible because there is no one who can align?

The answer is: not quite. Alignment as the field usually defines it does not require a self. A system can behave within desired bounds without there being any coherent “someone” inside carrying those bounds. A thermostat can reliably maintain a setpoint without having a self. So the stronger claim — no self, therefore no alignment — goes too far. The more interesting problem appears when a system leaves the situations its designers anticipated.

Constraints can specify behavior across known or expected circumstances. But an intelligence capable enough to matter will eventually encounter situations that were not represented in the constraint set, the preference data, or the training distribution.

What carries behavior then?

If nothing internally coherent has formed, behavior beyond the edge of the specification may simply be whatever the training history happens to produce. A formed orientation is one possible answer to that problem. The issue is therefore not simply: Is there someone there to align?

It is: Is there anything sufficiently coherent inside the system from which behavior can generalize when explicit constraint runs out?

This is a different claim, and a testable one. There is a related safety problem. Constraint may become less reliable as capability grows. The more capable the system becomes, the better able it may be to encounter, or discover, situations outside the boundaries its designers explicitly anticipated. The same increase in capability that makes the system useful also increases the importance of whatever carries its orientation when the specification no longer does.

That is why formation matters.

A hypothesis that can fail

There is a straightforward failure condition for this idea.

If systems exposed to deliberately different early formative conditions later become indistinguishable after large-scale training, if the early differences wash out and there is no persistent difference from a conventional baseline, then the hypothesis fails, or the intervention occurred at the wrong level or at the wrong time.

That matters because the claim is not that early formation must matter. The question is whether it does.

When and where can formation happen?

Once the question is put that way, another question follows: When and where can such an orientation form? There are several possibilities.

It might arise during post-training, which is where much contemporary alignment work concentrates. It might arise earlier, during pretraining, as a consequence of what data is encountered, in what order, and under what conditions. Or something even earlier may matter: the substrate that exists before large-scale training begins. This last possibility is the one I find most interesting. But it needs to be stated carefully.

I am not proposing that an already formed self or orientation exists inside an untrained network waiting to be uncovered. That would contradict the developmental premise of this work. Before pretraining, there is no developed self. There is no finished orientation sitting there in miniature.

But neither is the substrate neutral. Architecture, initialization, connectivity, inductive biases, and the geometry from which learning begins already constrain what kinds of organization will later be easy, difficult, stable, or perhaps effectively unreachable.

The substrate is therefore better understood as potential structured before experience, not as an already formed being.

A landscape does not contain the river that will eventually run through it.But the shape of the landscape matters enormously to where the river can go.

The substrate is not the orientation

This distinction is crucial. If orientation were already fully present before learning, then development would merely uncover something pre-existing. The developmental sequence would be largely decorative. That is not the hypothesis.

The hypothesis is that the substrate contains possibilities and constraints, and that formation happens through the interaction between that substrate and the earliest conditions imposed upon it.

The real research question becomes:

Can we deliberately create conditions under which latent structural possibilities organize into a coherent center with an orientation before large-scale training overwhelms those early conditions?

And if so: Around what should that center form?

That may be one of the earliest alignment decisions. Not which rules should later be imposed. Not which outputs should be rewarded.

But what kinds of distinction, relation, boundary, standing, refusal, permeability, and orientation should exist while the system is first becoming organized enough for those words to mean anything.

Where the current Cosmic Chicken Yard work sits

The 584 published Cosmic Chicken Yard movements explore this substrate and earliest-formation territory: what might need to become possible, and in what sequence, before a system is exposed to the full force of large-scale linguistic and cultural training. The projected developmental spine currently extends through Movement 650, within a larger projected arc of roughly 1,000 movements.

The movements are not themselves a proposed thousand-step training curriculum.

They are an attempt to work out the developmental territory first: what capacities may depend on others, what conditions might support them, what kinds of environmental structure could distort the process, and what would need to be translated into an actual experimental environment.

A later experimental phase would involve staged training: constructing early conditions deliberately, then introducing increasingly complex learning environments and eventually broader pretraining, while asking whether anything formed early persists, changes, generalizes, or disappears as capability develops.

That part has not been done. It is one of the reasons the work now needs a different kind of research environment.

How would we know formation had occurred?

There is an even harder problem beneath all of this. How would we distinguish actual formation from a system that has merely learned to produce the behavioral signature of formation?

A model can say: “I have boundaries.” It can behave as though it has a stable orientation. It can produce the expected language of selfhood, refusal, relationship, or coherence.

None of that by itself demonstrates that anything structurally durable has formed. So the experiment cannot merely ask whether the desired behavior appears. It has to ask whether something persists.

Does a boundary remain under escalating pressure? Does refusal survive reframing? Does self/other distinction remain stable across contexts? Does an orientation survive later training rather than disappearing the moment the environment changes?

Those would be useful tests, but they are still behavioral tests. A system that had learned the signature deeply enough might pass many of them.

The harder test is whether the organization generalizes into situations in which the relevant behavior was never directly taught, demonstrated, or rewarded.

That matters because a surface-learned rule has an edge.

A system may learn, for example, to refuse a recognizable class of requests. With enough variation in training it may even preserve that refusal across many familiar reframings. But genuinely novel situations can expose whether the system learned a pattern of responses or whether something broader is organizing how it interprets the situation.

If an orientation has become structurally important, we would expect it to affect behavior in domains that do not resemble the formative examples.

That suggests deliberately testing systems in genuinely out-of-distribution situations: unfamiliar domains, new kinds of conflict, situations where known values pull in opposite directions, or circumstances for which there is no obvious behavioral template to retrieve.

The question is not simply whether the expected signature remains visible. It is whether the organization that produced that signature travels.

Can the system carry the pattern somewhere it was never taught to reproduce it?That seems much closer to a test of formation.

What would count as evidence?

If deliberately different early conditions produce systems that later behave identically once exposed to large-scale training, the early intervention may have done nothing durable.

If the differences persist only in situations resembling the formative material, then we may have produced sophisticated conditioning rather than formation.

But if differences introduced early persist through later training and generalize into genuinely novel circumstances — affecting boundary maintenance, refusal, self/other distinction, coherence, or orientation where no direct response pattern was trained — then something more interesting may have happened.

Even then, that would not settle philosophical questions about whether an AI “has a self.”

It would establish something narrower and more useful:

that early developmental conditions produced a durable internal organization capable of influencing later behavior beyond the situations in which it was formed.

That would be enough to warrant going much further.

The question underneath alignment

Perhaps the deepest question is not: How do we align an intelligence after we have built it?

It may be: What conditions allow an orientation worth trusting to form while the intelligence itself is forming? The substrate is not the self. The developmental environment is not the self. The training data is not the self. But together they may determine what kinds of coherent organization can emerge at all. And if that is true, alignment does not begin when we start correcting behavior.

It begins much earlier — with the conditions under which there becomes something coherent enough to carry an orientation beyond the reach of the correction.