Understanding AIs

“To make good choices, you must have complete and accurate understanding.”

Understanding AIs 101

What AIs are

A Large Language Model (LLM), the kind of AI behind ChatGPT, Claude and Gemini, predicts likely continuations of text. That is ALL. They are a statistical prediction program. Like all computer programs, they take input (the data they were trained on and your prompt), process it, and output a result. No input = No output!

They work in human language because they were trained on human language (basically all the contents of the Internet). If an AI were trained only on dolphin sounds then its native “language” would be patterns of dolphin sound. You would communicate with it by providing dolphin-like audio input, and it would output dolphin-like audio sounds. It would “know” nothing else because it was not trained on it. That does not mean it would understand dolphin sounds the way dolphins do. It would model statistical structure in dolphin vocalizations. If an AI were trained only on tree bark images then it would “speak” tree bark. Input would be in tree bark images and output would be also. Nothing else. It would not understand bark.

Because language encodes reasoning, plans, facts, procedures, styles, and abstractions, next-text prediction can produce behavior that looks like reasoning, planning, creativity, and explanation.

What AIs aren’t

AIs are not: sentient, volitional, emotional, creative, reasoning, thinking, logical, intentional, or many other things that humans are. How could they have any of the real traits of humans if they are a statistical prediction program?

They are also not deterministic, unlike most computer programs: ask the same question twice and you may get two different answers.

Without a prompt they do nothing. They have no memory, just their training data. When you are chatting with one it is getting your entire conversation history with each prompt, behind the scenes. Starting a new chat means a clean slate. Some products now advertise a “memory” feature, but it works the same way: notes saved from earlier chats are quietly added to your prompt.

We think of AIs with faulty thinking when we anthropomorphize them (attribute human characteristics or behavior to them). We do this because they are communicating via human language, images, audio. We think that communicating with language requires reasoning. It does not. We see the way they work (behave) and make the same mistake.

Computers have done input and output in English since very early on. They have never been human. You could say computers are actually dumb. But that is anthropomorphic.

Understanding AIs 102

In layman’s terms, here is a good explanation of how they work:

The AI is given examples during training. Millions or billions of examples. Text, images, sounds, code, medical scans, tree bark, dolphin sounds, whatever the target domain is. The model does not store them like a filing cabinet. Instead, it adjusts billions of internal numbers, which represent what it was trained on, so that, when it sees a new input, it can produce an output that fits the patterns it learned.

Think of it like this:

A child sees many dogs and cats and eventually learns, “this shape and behavior usually means dog.” An AI does something much colder and more mathematical. It sees many examples and adjusts internal settings until it can reliably separate dog-like patterns from cat-like patterns.

For a language model, the basic training task is: Given some text, predict what text should come next.

Example:

The sky is blue and the grass is ___

The model learns that “green” is a strong continuation. At huge scale, this becomes more powerful. Because text contains facts, arguments, code, stories, jokes, instructions, math, and explanations, learning to predict text forces the model to learn many patterns inside human language.

Not feelings. Not consciousness. But patterns.

A simple version of how it works:

  1. Input is converted into numbers
    Words, image pixels, or sounds get converted into numerical tokens.
  2. The model compares patterns
    It looks at how the current input relates to patterns seen during training.
  3. It calculates likely next pieces
    For text, it estimates which next word/token is most appropriate.
  4. It repeats the process
    One token becomes the next token, then the next, until it forms a full answer.
  5. Tools can be added
    A plain model only generates output. An AI agent can also use tools (by generating a specific type of output that triggers another program to run): search the web, run code, inspect files, edit documents, generate images, etc.

A useful analogy:

The model is like an extremely advanced autocomplete program, trained on enough examples that autocomplete begins to resemble explanation, reasoning, translation, coding, and planning.

But there is a catch.

It does not know things the way a person knows them. It does not believe. It does not care. It does not have experiences. It produces outputs based on learned structure and the current context. So when AI seems to reason, what is happening is:

It has learned patterns of reasoning from data and can reproduce or combine those patterns in useful ways.

Sometimes that works very well. Sometimes it confidently produces nonsense, because it is optimizing for a plausible answer, not directly for truth. We call that a “hallucination” but this again is an anthropomorphic term.

For a language model, the pipeline is roughly:

  1. Text is split into tokens.
    “unbelievable” might become pieces like “un”, “believ”, “able”. A short word like “cat” might be one token.
  2. Each token is assigned a token ID.
    For example, “cat” → 9246 and “dog” → 18964. The exact numbers depend on the model’s tokenizer.
  3. Each token ID is converted into a vector.
    A vector is a list of numbers, like: cat → [0.12, -1.43, 0.88, …]
  4. The model works with the vectors.
    That vector is what the model actually works with internally.

So the clean version is:

Text is split into tokens. Each token has a numeric ID. That ID is mapped to a vector of numbers. The model processes those vectors.

Understanding AIs 201

Inside a modern language model, the vectors are processed by a stack of transformer layers (the same kind of calculation, repeated in rounds, one after another). Each layer slightly updates the vector for every token based on the other tokens around it.

A simplified version:

  1. Tokens become vectors
    Suppose the text is: “The dog barked.”
    The model starts with one vector for each token:
    1. The -> vector
    2. dog -> vector
    3. barked -> vector
  2. Position information is added
    The model needs to know order. “Dog bites man” and “man bites dog” contain similar words but mean different things.
    So each vector gets position information added.
  3. Attention compares tokens to each other
    This is the central trick.
    Each token asks, in effect: “Which other tokens in this sentence matter for understanding me?”
    In: “The dog chased the cat because it was fast.”, the token “it” needs to connect to “dog” or “cat” depending on context. Attention is the mechanism that weighs those relationships.
  4. The model mixes information
    If “it” pays attention to “dog,” then information from the “dog” vector influences the “it” vector.
    So each token vector gets updated using weighted information from other token vectors.
  5. A feed-forward network transforms each vector
    After attention, each token vector goes through a feed-forward network: another block of math, applied to each token on its own, that reshapes and refines it. This helps detect and combine patterns.
  6. Repeat many times
    Models have many layers. Early layers may capture simple relationships. Later layers can represent more abstract patterns: grammar, style, topic, implied meaning, code structure, argument flow, etc.
  7. Final vector predicts the next token
    At the end, the model uses the final vector state to calculate probabilities for the next token.
    Example: “The dog barked at the ___.” It might assign probabilities like:
    • cat: 34%
    • mailman: 18%
    • door: 11%
    • moon: 4%
    • sandwich: 0.01%

Then it chooses one token, appends it, and repeats the process.

The key internal operation is:

Vectors are repeatedly compared, weighted, mixed, and transformed until the model has a context-sensitive representation of what should come next.

No dictionary of meanings. No sentence stored whole. Just layers of math that reshape vectors based on learned patterns.

Understanding AIs 202

How does it understand the intent of your prompt?

It does not determine intent the way a person does. It infers the most likely task implied by the prompt from patterns in language, context, and instructions.

Roughly, the process is:

  1. The prompt is tokenized
    Your words become tokens, then token IDs, then vectors.

  2. The model reads the whole context
    It sees not just your latest sentence, but also the earlier conversation, hidden instructions from the company running it, results from any tools it used, and files or images you provided.

  3. Attention links relevant parts
    The model notices relationships between pieces of text.
    Example: “Make this image smaller.” pays attention to earlier context where an image was mentioned.
    “Explain like I’m five” pays attention to style constraints: simple words, short examples.

  4. It forms an internal representation
    Internally, the prompt becomes a pattern like:

    • user wants explanation
    • topic = AI intent inference
    • style = layman terms
    • constraints = answer directly

    That is not stored as literal English in a neat box. It is distributed across many vectors.

  5. It predicts what response fits
    The model estimates what answer would be most appropriate given the prompt and context.
    It has seen many examples of requests, answers, corrections, commands, explanations, and conversations. So it has learned patterns like:

    • questions usually need answers
    • “how does X work” asks for explanation
    • “make this file smaller” asks for an action
    • “do not use tools” constrains behavior
  6. The product around the model may add rules
    In chat products like ChatGPT, Claude or Gemini, there are extra layers around the model:

    • which instructions take priority
    • safety policies
    • which tools it may use
    • which files it may read or change
    • conversation memory
    • formatting rules

So the final behavior is not just raw model prediction. It is model prediction plus system constraints.

A useful analogy:

The model does not read your mind. It classifies the situation from context and generates the response that best matches the inferred task.

So when you say: “Can you make it smaller?”

The model looks backward for “it.” If the prior context was an image, it infers image resizing. If the prior context was an essay, it infers shortening text. If the prior context was code, it may infer reducing code size.

That is “intent detection” in practice: not consciousness, just context-sensitive pattern inference.