Chip Joyce
Essays

Creativity, Conjecture, and AI

Epistemology

When Sergey Rachmaninoff sat down to compose his Piano Concerto No. 2, he was not extrapolating from existing concertos. He was not averaging Tchaikovsky and Chopin to find some optimal middle ground. He was recovering from a creative crisis, haunted by the failure of his First Symphony, working with a hypnotherapist to rebuild his confidence. What emerged captured both melancholy and triumph in a way no prior composition had.

That was not pattern-matching scaled up. It was conjecture.

Most of us misunderstand how we learn and create. We imagine knowledge entering the mind from experience, like water filling a vessel. We think creativity means recombining existing elements in novel ways. We assume that with enough data and processing power, any system could eventually do what Rachmaninoff did, what Einstein did, what the SpaceX engineers did when they worked out how to land rockets vertically.

None of that is how it works. The real mechanism is the thing our machines do not have.

How Anyone Learns Anything New

Here is the puzzle. If knowledge comes from experience, we are stuck. Experience can only show what has already happened. It can reveal patterns in existing data. It cannot tell us what lies beyond that data, explain why the patterns exist, or say whether they will continue. Every scientific law, every creative breakthrough, every solution to a hard problem requires going beyond what experience alone can justify.

David Deutsch puts it plainly: we do not learn by extracting patterns from data. We learn by guessing—by creating explanatory conjectures that reach beyond our observations, then testing whether reality contradicts them.

Conjecture is the leap a mind makes when it proposes an explanation or a solution not contained in prior experience. It is guessing, but not random guessing. It is disciplined imagination constrained by what you are trying to explain.

When Einstein wondered what it would be like to ride alongside a beam of light, he was not interpolating between existing theories. He was imagining a scenario that violated common sense and working out what must be true if it made sense. Special relativity came from that conjecture, not from more data about moving objects.

When SpaceX engineers proposed landing orbital-class rockets vertically, they were conjecturing against the entire history of spaceflight. Every prior rocket had been expendable. The data suggested, strongly, that expendable was correct. The conjecture was: what if we are wrong about what is possible? They built an explanatory framework for why controlled descent should work, then tested it against reality. Seventeen landing attempts failed before the first success. The pattern in the data was failure. The conjecture saw past it.

They were not finding patterns in existing rocket-landing data. There was no successful data to match against. They imagined a solution, modeled it, worked through the physics, anticipated failure modes, and proposed fixes before those failures occurred. Each iteration involved new conjectures.

This is how humans solve hard problems. We imagine explanations. We do not extract them from data.

What the Machine Does Instead

A large language model finds statistical patterns in its training data. It identifies which tokens tend to follow which other tokens in which contexts, and given a prompt, predicts the most likely continuation. It is extraordinarily sophisticated pattern-matching—billions of parameters capturing unimaginably complex statistical relationships—and it is still pattern-matching.

One can debate whether human cognition is itself pattern-matching at the biological level. Set that aside. Even granting it, the difference in output is real: a model cannot imagine something outside its training distribution. It can interpolate impressively within that distribution and combine patterns in ways that seem creative. It cannot make the leap Einstein made, or the one those engineers made.

Ask a model to solve a novel problem and it finds the most statistically likely response based on similar problems it has seen. When a human solves a novel problem, he creates an explanation for why a particular approach should work, and that explanation may have no precedent in his experience.

Consider the discovery of Neptune. In the 1840s astronomers noticed that Uranus was not moving as Newton’s laws predicted. There was a pattern in the observations: Uranus kept deviating from its expected orbit.

A pattern-finding system would have described that deviation precisely. It might have projected the pattern forward and predicted further deviations.

Instead, astronomers conjectured an unseen planet whose gravity was pulling on Uranus. They calculated where it would have to be to produce the observed deviations. They pointed their telescopes at an empty patch of sky.

And there was Neptune.

The data alone could never have produced that. The data showed deviations. It took a mind to imagine an invisible planet and then test whether the guess was right.

Language models are useful for a great many things. They summarize, suggest code, draft routine communications, answer factual questions where the answer exists in their training data. These are real capabilities.

They cannot imagine Neptune.

The Intent Behind the Notes

Return to Rachmaninoff. He had absorbed the tradition—all artists work within traditions. Then his mind did something more specific than novel recombination. It formed an intent. He was trying to express something particular, a feeling with no prior musical form, an emotional experience that needed exactly these notes and not others. That intent shaped every decision. Each phrase was a conjecture tested against a single question: does this convey what I am reaching for? He would play a passage, listen, judge whether it achieved what he intended, and revise.

The concerto emerged through thousands of small conjectures, each anchored to a target that existed only in his mind before it existed in the music.

A model trained on classical music could generate something that sounds like Rachmaninoff. It could find the patterns in his style and produce music statistically similar to his compositions. Some listeners would not immediately hear the difference.

Statistical similarity is not what Rachmaninoff was doing. He was not producing music in his style. He was expressing something specific, something he had imagined, something that required the notes he wrote and no others. A pattern-matching system has no target it is reaching toward, only likelihoods it is satisfying.

Scaling Does Not Bridge This

Language models will get larger, more sophisticated, more capable of finding complex patterns in massive datasets. They will become more useful for more tasks.

Scaling pattern-matching does not eventually produce conjecture. These are different processes, not different points on one spectrum. A system that finds patterns in data, however sophisticated the patterns or vast the data, is doing something different from a mind that creates explanations reaching beyond any data. More parameters do not bridge that. More training data does not bridge it. The difference is in kind.

When people claim we are close to artificial general intelligence, they are not seeing this. They are impressed by how well models pattern-match within their training distribution, and they assume that more of the same eventually produces what a mind does. That is like assuming a more powerful calculator will eventually become conscious.

Use these tools where they excel: analyzing data, finding patterns, extending human capability in specific domains. Stop pretending we are building minds.

Marvel instead at the minds we already have. Your ability to read this and imagine counterarguments, to conjecture whether any of it is right, to think of examples I have not mentioned that might support or contradict it—that is the extraordinary thing.

The astronomer imagining Neptune in an empty patch of sky. Rachmaninoff imagining music that did not exist. Engineers imagining reusable rockets when all the data said it could not work. You, right now, imagining whether this holds.

That is what the machines cannot do.