The smarter the models get, the clearer it becomes: what holds us back is not their intelligence but the way we communicate with them.

A familiar situation

You ask an AI to do something and get almost what you wanted. You explain what is off. Then you explain again. By the tenth revision you realize that you know exactly what you want, you can practically see it. You just cannot say it.

At this point people usually decide the model is not smart enough. We think the problem is somewhere else.

Two ways the mind works

We separate two modes of thought.

Word-thinking works with words. It is good where the answer is already known: explaining, checking, passing it on to someone else.

Image-thinking works with images. An image here is not a picture. It is a whole sense of a situation in many dimensions at once: shape, weight, motion, time, balance. To think this way is to move that image around and watch what happens to it.

Imagine someone hands you an antique tool you have never seen. You have no name for it. But your hands are already turning it over, weighing it, finding how it sits in your palm. A minute later you are holding it correctly, though you still cannot explain why. Understanding arrived before the words did.

We find new things through image-thinking. Words come afterwards, to keep what we found and tell others about it.

Words as a bottleneck

To share a thought you have to compress it into words. The other person unpacks those words into an image of their own. Something is lost twice, as anyone knows who has tried to describe a melody in writing instead of humming it.

The same thing happens with AI, and here is the curious part. There are no words inside a language model either. Inside is a vast space of numbers where meanings sit closer to or farther from one another, smoothly. Words appear only at the input and the output.

So the picture is a strange one. On both sides of the conversation there is something rich and continuous, and between them a slot one line of text wide.

What if we moved it instead of describing it

We imagine a different way of working. The model shows you an image. You feel that something is off and you shift it with a glance, a gesture, a small movement. The model responds at once, and you keep moving until it feels right.

It is the difference between turning a volume knob and writing a letter that says “please make it twelve percent quieter.”

None of this requires a chip in your head. A first version of such an interface can be built from a screen, a camera and models that already exist. A direct link to the brain is a distant prospect, and we are not promising it tomorrow.

How we want to test this

The experiment is called the Flashlight Beam. A person in a dim room is given an unfamiliar antique object. Light falls only on their hands and on the object. The task is to work out what it was made for and show it through action.

Cameras record every movement of the hands and eyes. Our prediction is simple: the hands will find the right grip before the person can explain the object's purpose in words. And if we ask them to reason out loud, they will do worse.

If the experiment does not show this, we are wrong. We will say so.

Our bets

We cannot prove these yet, but this is where we are placing our bets.

  • Text chat with AI is a temporary form. In ten years it will look the way the command line looks today.
  • The next big step will come not from a smarter model but from a wider channel between the model and the person.
  • Image-thinking can be taught. It takes a master who shows you what you just did, and that master can be an AI.
  • AI is more useful as a mirror than as an oracle. A machine that helps you see the course of your own thought will give you more than a machine with ready answers.

Argue with us

These are hypotheses, and we are publishing them to start a conversation. The question we find most interesting is this: is there a problem that cannot be solved without words?

If you have an answer or an objection, write to us.