The | Gen AI | AI Agents | AI Tools | Quantum | QAI | QML | Research | Developer | Python | Golang | Rust | Mojo | Mastery Playbook.
The | Gen AI | AI Agents | AI Tools | Quantum | QAI | QML | Research | Developer | Python | Golang | Rust | Mojo | Mastery Playbook.
The | Gen AI | AI Agents | AI Tools | Quantum | QAI | QML | Research | Developer | Python | Golang | Rust | Mojo | Mastery Playbook.

In 2006, I was in my first year after graduation.
I was working on Hopfield networks in C.
I was also learning Non-Linear Dynamical Systems theory at the time.
I suspected very strongly that neural networks, being dynamical systems, had chaotic boundaries and strange attractors of their own.
My research ambition was to find them.
However, hardware at that time did not allow me to run thousands of Hopfield Networks in parallel.
However, it was an interesting theoretical question, and it intrigued me.
Now – stuff happens – life happens – and I took 4 years to get a two-year degree, because of my poor health at the time.
I did not think that CS was my field, and went into theology for 5 years.
The worst decision of my life – without qualification.
However, one thing was clear – my place was research, and I found interesting connections between quantum mechanics and theology(!) – but that’s an article for another time.
I slowly moved into technical writing, the one place where I could still research and form theories of my own.
Then transformers came out.
Deja-vu!
I realized transformers were non-linear dynamical systems, and wanted to run billions of them as well!
Life – and other things – had other ideas.
Imagine my delight when I revisited the topic in ten years later, 2026, and found that more scientists had come to similar conclusions.
This article is an interesting adventure.
I am going to feed you the key concepts in nonlinear dynamical systems theory through images and their history.
All the images in this article are deep.
Really deep.
And I hope you come to see the beauty in them that I see in them.
I am 40 years old and no longer in college.
My dreams of research have been replaced with Quantum Computing and Generative AI – which is also cool – because they too are ripe fields for research.
So, over to AI, because, finally, through chaos and complexity theory – we finally have a way to start understanding the behaviour of these incredibly complex systems.
Ask an honest AI researcher how large language models work, and you will get an uncomfortable answer.
"We know how to build them. We do not fully know how they work."
That is not false modesty.
It is the actual state of the field.
We can write down the architecture.
We can write down the training loop.
We can tell you every matrix multiplication in exhaustive detail.
And yet nobody can tell you, from first principles, why a model at 15 billion parameters can do arithmetic when a model at 10 billion cannot.
Why does "think step by step" work?
Why do capabilities appear in sudden jumps rather than smooth climbs?
Why does a model that has clearly memorized its training data suddenly, thousands of steps later, start actually understanding?
I'm writing this article to tell you that these questions are not unanswerable.
They only look unanswerable because most people are looking at LLMs through the wrong lens.
Stop looking at a language model as a very complicated program.
Start looking at it as a nonlinear dynamical system.
Do that, and the fog lifts.

The sudden jumps become phase transitions – the same mathematics that governs water turning to ice.

The mysterious plateaus become competitions between internal circuits.

The sensitivity of long reasoning chains becomes the butterfly effect.

The architecture tricks that "just work" turn out to be criticality-engineering devices borrowed, unknowingly, from the physics of sandpiles and earthquakes.

This is published, measurable, verifiable  science – and it comes with real numbers.
Read on!
Let's define the term, because everything in this article depends on it.
A dynamical system is anything that changes over time according to rules.
The weather. A pendulum. A stock market. A heart. A population of rabbits.
Linear means effects are proportional to causes.
Push twice as hard, go twice as far.
Double the input, double the output.
Linear systems are predictable, well-behaved, and – frankly – rare in nature.
Nonlinear means the system feeds back on itself.
Output becomes input.
Small causes can produce enormous effects, and large causes can produce nothing at all.
The response is not proportional to the push.
That single property – feedback – is what makes reality interesting.
And nonlinear dynamics splits into two great research traditions, which are the two halves of this article:
Chaos theory studies how simple deterministic rules produce unpredictable behaviour. There is no randomness in a chaotic system. It is fully determined. But it is so exquisitely sensitive to its starting conditions that prediction becomes practically impossible. That is why we can know the exact physics of the atmosphere and still not forecast rain three weeks out.
Complexity theory studies the opposite miracle – how vast numbers of simple parts, interacting locally, spontaneously produce sophisticated global order that nobody designed. Ant colonies. Immune systems. Cities. Brains.
Chaos is order producing apparent randomness.
Complexity is randomness producing apparent order.
An LLM does both. At once. At enormous scale.
Now here's your first pair of terms.
ChatGPT user:
AI engineer:
Hold onto that.
We will need it.

The foundational fact of LLM science is the scaling law.
In 2020, OpenAI published Scaling Laws for Neural Language Models, showing that model quality improves in a smooth, predictable curve as you add parameters, data, and computing power.
Parameters are the adjustable dials inside the model.
GPT-3 had 175 billion. Each one gets tuned during training.
More dials means more capacity to store patterns.
Loss is a score for how badly the model predicts the next word.
Lower is better.
It is the only thing training actually optimizes – everything else is a side effect.
Tokens are chunks of text, roughly three-quarters of a word each.
A 300-page book is about 100,000 tokens.
The equation, refined by DeepMind's Chinchilla paper, is:
Loss = E + A/N^α + B/D^β
So:
In English:
That is the whole industry in one line.
And notice its shape.
It is a power law – and power laws are the fingerprint of nonlinear systems everywhere in nature.
ChatGPT user:
AI engineer:
DeepMind trained over 400 models to find the right balance.
The answer: roughly 20 tokens of training text per parameter.
A 70-billion-parameter model wants around 1.4 trillion tokens.
Why does this matter?
Because before Chinchilla, everyone got it wrong.
GPT-3 used only about 1.7 tokens per parameter – it was starved of data.
Chinchilla-70B beat the 280-billion-parameter Gopher while using four times less computing power, purely by rebalancing.
ChatGPT user:
AI engineer:
Smooth curves. Predictable costs. Everybody happy.
And then the smooth curves broke!

In 2022, Google researchers published Emergent Abilities of Large Language Models and found something the scaling laws did not predict.
The loss improves smoothly.
But actual skills do not.
Skills stay at zero.
Then they appear, almost fully formed.
Ability
Appears around
Basic zero-shot task following
~1.5 billion parameters
Two and three-digit arithmetic
~13 billion parameters
In-context learning (learning from prompt examples)
~2.5-5 billion training tokens
Chain-of-thought reasoning becoming useful
~60-100 billion parameters
Complex multi-step logical reasoning
~540 billion parameters (PaLM era)
This is the language of physics, and physics has precise words for it.
ChatGPT user:
AI engineer:
Look closely at row three of that table.
That threshold is not measured in model size at all – it is measured in how much text the model has read so far.
We will come back to it.
It is the most beautiful result in the field.
ChatGPT user:
AI engineer:
Fair challenge, and a serious one.
In 2023, Stanford researchers published Are Emergent Abilities of Large Language Models a Mirage?.
Their argument: if you grade with a harsh all-or-nothing metric, you manufacture fake cliffs.
A model getting steadily better at arithmetic looks like zero right up until it is perfect.
They were partly right.
Some emergent abilities do melt away under gentler grading.
But not all of them.
Sharp jumps survive even under smooth continuous measurements – and the biggest transition of all is visible directly in the loss curve, which is about as continuous as measurement gets.
The critique did not kill the field. It made it rigorous.
That is exactly how science is supposed to work!

Here is the insight I find most elegant in all of this.
The thresholds above are written in parameters.
But parameters are not what causes the jump.
Pretraining loss is.
Think of it this way.
Loss is the temperature.
Parameters and data are just the stove settings you use to get there.
Water freezes at 0°C whether you cooled it in a freezer, a cold room, or a mountain cave.
The route does not matter.
The temperature does.
So an ability does not "appear at 13 billion parameters."
It appears when prediction error drops below a critical value – and 13 billion parameters happened to be how you got there in 2022.
ChatGPT user:
AI engineer:
And this explains something that puzzles everyone.
ChatGPT user:

Now for the strangest experiment in machine learning.
In 2022, researchers trained tiny models on modular arithmetic – clock math, where 9 + 5 = 2 on a 12-hour clock.
They published Grokking: Generalization Beyond Overfitting.
The model memorized the training examples fast.
Perfect scores on questions it had seen.
Zero ability on new ones. Classic cramming.
Then the researchers kept training. And training.
For thousands more rounds while absolutely nothing changed.
Then – suddenly – the model generalized perfectly!
It had stopped memorizing and started actually understanding.
Memorize.
Plateau.
Snap.
Understand.
Because two strategies compete inside the network.
Researchers call it circuit competition.
Strategy A – the lookup table.
Memorize every answer. Fast to learn, but bulky. It eats capacity.
Strategy B – the algorithm.
Actually work out the rule. Slow to discover, but tiny and elegant – in this case, the model literally invents a Fourier-transform-based method.
Training applies gentle pressure toward simplicity.
Early on, memorization wins because it is quicker.
But as training grinds on, the elegant algorithm – far cheaper to maintain – eventually outcompetes and replaces it.
The plateau is not stagnation.
It is a war being fought inside the weights.
ChatGPT user:
AI engineer:
And the theory made falsifiable predictions that came true – including ungrokking, where a model that already generalized reverts to memorizing if you shrink its training set.
When a theory predicts a bizarre new phenomenon and experiment confirms it, that theory is doing real work.
A 2024 unification paper then showed this single mechanism explains three separate mysteries at once: grokking, the double-descent curve, and emergent abilities at scale.
Three puzzles.
One explanation.
ChatGPT user:
Generalization is compression.
Occam's Razor, enforced by gradient descent.

Time for chaos theory proper.
Imagine whispering into one end of a very long chain of people.
Two possible failures:
Too much damping – each person mumbles slightly quieter, and by person fifty the message is silence. Information dies. The ordered phase.
Too much amplification – each person exaggerates slightly, and by person fifty you have gibberish. Small differences explode. The chaotic phase.
Neither chain can carry a message.
But there is a razor-thin setting in between – the edge of chaos – where signals travel any distance without dying or exploding.
Deep neural networks are exactly this chain, with layers instead of people.
And a landmark NeurIPS 2004 paper established that only networks near this critical boundary can perform genuinely complex computation.
Later work made it mathematical.
So:There is a precise critical line in the settings used to initialize a network's weights, and information survives arbitrarily deep only exactly on that line.
ChatGPT user:
AI engineer:
Here is the part that should make you sit up.
Every architectural trick in the modern transformer – LayerNorm, residual connections, careful initialization – exists to hold the network at this edge.
Residual connections give signals a clean path through.
Normalization acts like cruise control, constantly nudging the network back to critical.
Without them, deep transformers suffer rank collapse, where every token's representation converges to the same thing – the ordered phase's information death, in production.
We did not design these components as chaos-theory devices.
But that is exactly what they are.
ChatGPT user:
AI engineer:
Every trick in the modern transformer playbook is, secretly, Per Bak's sandpile wearing a lab coat.
How cool is that?
And it explains the temperature slider you use every day.
ChatGPT user:

I saved this one for last.
Anthropic researchers studied 34 transformer models during training and published In-Context Learning and Induction Heads.
In-context learning is the model's ability to learn from examples inside your prompt.
Show it three examples of a format, and it follows that format on the fourth.
No retraining.
This is the single ability that makes prompt engineering possible at all.
They found that every transformer deeper than one layer undergoes an abrupt internal transition at roughly 2.5 to 5 billion training tokens.
Before that point: essentially no in-context learning.
After it: in-context learning simply exists.
The circuits responsible are called induction heads – attention patterns that spot "this happened before, here is what followed" and complete the pattern.
And here is the astonishing part.
This threshold barely depends on model size.
A 100-million-parameter model crosses it at roughly the same token count as a 500-billion-parameter one.
It is a time-based transition, not a size-based one.
You can see the bump in the loss curve with your naked eye.
It is the closest thing our field has to watching ice crystallize.
ChatGPT user:
AI engineer:
Every time few-shot prompting works for you, you are using a circuit that formed in a single abrupt moment during that model's training.

Print this and pin it up:
The size numbers keep dropping.
The loss thresholds do not.

So where does this leave us?
It leaves us with a field that has barely started.
Treating an LLM as a nonlinear dynamical system is not a cute reframing.
It is a research programme, and it opens doors in every direction.
Capability forecasting.
If skills are loss-threshold phenomena, we can predict what a model will do before we finish training it.
That is enormous – scientifically and commercially.
Safety.
Anthropic researchers warned in Predictability and Surprise in Large Generative Models that dangerous capabilities may appear abruptly rather than gradually.
Physics has spent a century learning to detect approaching phase transitions before they happen – divergent susceptibility, critical slowing down, growing correlation lengths.
Every one of those tools is now being pointed at training runs.
Early-warning systems for capability jumps are being built right now.
Better training.
Understanding grokking already produced Grokfast, an optimizer that deliberately accelerates the memorization-to-understanding transition.
That is theory paying rent within two years.
Architecture design.
If depth requires criticality, then designing for criticality directly – rather than stumbling into it through LayerNorm and residuals – is an obvious frontier.
Four avenues.
And they are only the ones visible from here.
But I want to close with the biggest one of all.
Because there is another nonlinear dynamical system that we have been trying to understand for far longer than we have had transformers.
It weighs about 1.4 kilograms.
It runs on twenty watts.
It sits behind your eyes as you read this sentence.
Your brain.
The human brain is the most complex nonlinear dynamical system we know of in the universe.
And – remarkably – real neuroscience has found the same signatures we have been discussing all article long.
Neural avalanches following power laws.
Cortical dynamics poised at criticality.
Phase transitions between states of consciousness.
Self-organized tuning to the edge of chaos.
The same mathematics.
The same fingerprints.
I am not claiming an LLM is a brain. It is not, and anyone who tells you otherwise is selling you something that is not rigorous science.
But I strongly believe we may have accidentally built the first nonlinear system complex enough to be interesting, simple enough to be studied, and – crucially – fully instrumented.
We can read every weight.
We can checkpoint every step.
We can rerun the whole thing with one variable changed.
You cannot do that with a brain.
So here is the possibility that keeps me up at night, in the best possible way.
What if the science we develop to understand emergence in language models becomes the science that finally cracks emergence in minds?
What if the tools we are building to detect a phase transition in a training run turn out to be the tools that explain how understanding arises in a child?
What if the deepest payoff of artificial intelligence is not the artificial part at all – but what it teaches us about the intelligence we already had?
Understanding.
Consciousness.
Meaning.
The image of God in the machinery of the mind.
I do not know if we will ever get there.
But I strongly believe the road runs through nonlinear dynamical systems, and I strongly believe it is the most fascinating road in science today – the road to understanding AI, and maybe even I, myself/itself.
The edge of chaos is where computation lives.
It may also be where meaning lives.
And we have only just started waking up to the possibilities.

This has been one of the most fulfilling articles I ever wrote (in partnership with AI, of course).
Nonlinear dynamical systems is not higher math.
I’m also neck-deep in Quantum Mechanics, so yes – I have a different point of view – even if that is physics.
It’s just math.
The language of the universe.
I strongly urge my readers to stop saying that we do not understand how LLMs work.
Chaos and complexity theory are excellent starting points.
For nonlinear dynamical systems theory, I recommend the following courses as a good starting point:
(They require a basic knowledge of STEM – which I believe most HackerNoon readers have)
1. Complexity Explorer — Introduction to Dynamical Systems and Chaos (Winter 2015)
Archived SFI MOOC led by David Feldman, Professor of Physics and Mathematics at College of the Atlantic. Ten units span iterated functions, differential equations, chaos and the butterfly effect, bifurcations, the logistic map, universality, phase space, strange attractors, and pattern formation. Prerequisite: basic high school algebra. Ran January–March 2015; no longer in session.
2. YouTube playlist — Shane Ross, Nonlinear Dynamics and Chaos
Ross bills it as a free course for newcomers, delivered as short topical videos based on Strogatz's book. Lecture one is a historical overview; the series then tracks the textbook through flows, bifurcations, limit cycles, Lorenz, maps, and fractals. Engineering-flavoured, visual, modular.
3. ChaosBook.org — Predrag Cvitanović A free living textbook plus the courses built on it. It covers the same material as Georgia Tech's PHYS 7224 and is aimed at PhD students, postdocs, and advanced undergraduates, spanning topology of flows, Smale horseshoes, periodic orbits, symmetry, and transfer operators. Assumes linear algebra, ODEs, probability, statistical mechanics, and Python. Warning – the most advanced in this list.
4. Class Central — Nonlinear Dynamics (Liz Bradley, Complexity Explorer)
The YouTube mirror of SFI's intermediate course, running about nine hours forty minutes across maps, transients and attractors, bifurcations, return maps, flows, ODEs, fixed points, stability, manifolds, and strange attractors, plus ODE solvers from Euler methods to production-grade ones. Assumes college calculus, physics, and programming.
So:

And after 20 years of an ambition of research –
At least I managed to open some minds to what a nonlinear dynamical system is.
I’m grateful.
I have not published a paper.
But I enjoy putting out articles like these, educating the public (a very discerning public) on a 20-year-old theory that I had while in college.
And to be honest?
That feels real good.
Glory, praise, and honor be to God, the Supreme Author of everything that exists.
Peace!

Every image in this article was generated by Nano Banana Pro.
The first draft of this article (and the captions) were written with Claude Opus 5.

The Six-Day Mystery That Rewrote AI's Price List
The | Gen AI | AI Agents | AI Tools | Quantum | QAI | QML | Research | Developer | Python | Golang | Rust | Mojo | Mastery Playbook.