Transistors per Microchip Increased by More than 50 million Times in 50 Years

Superfact 127: The Intel 4004 processor chip created in 1971 had 2,308 transistors. Modern cutting-edge processors and AI accelerators exceed 50 billion to over 200 billion transistors, which is 25 million to nearly 100 million times as many transistors per chip. This follows Moore’s law, which states that the number of transistors on a microchip doubles roughly every two years, while the cost of computers drops. As a result, computer chips have gotten millions of times faster and the cost of memory (RAM) has dropped by four trillion times. This has made the modern AI / LLM possible.

It is a modern CPU on a white background. | Transistors per Microchip Increased by More than 50 million Times in 50 Years
CPU (a processor or a central processing unit). Concept of technological advancement. Shutterstock asset id: 2760284753 by pisanstock

The transistor below is a large transistor that you can hold between your fingers. Transistors inside a computer chip are very small. They are microscopic. In the diagram, whenever a small voltage ( typically 0.7 volts) is applied to the base (B) it allows a much larger current to flow through the main terminals from the collector (C) to the emitter (E). Since electrons are negative this means that the electrons are flowing from the emitter to the collector. If there is no voltage or current at the base there is no current flowing from the collector to the emitter either. This is how you control the transistor and achieve more complex behavior if you have many connected transistors. Also, note that the transistor is a one way street.

A large black transistor with its diagram. The diagram includes the base and the emitter and collector.
A large NPN-BJP transistor with its symbol diagram on the right. Shutterstock asset id: 2169223279 by Surkhab Ahmad Art

Moore’s law, which states that the number of transistors on a microchip doubles roughly every two years, is not a physical law, it is a purely observational law that has held for well over half a century. It is just a reflection of how good we are at improving processors. No one knows whether Moore’s law will continue to hold but if it does, then today’s computers and today’s AI are very bad compared to tomorrows computers and AI, and they are super bad compared to what we will have 20 years from now.

The diagram below shows 100+ microprocessors (CPUs), their number of processors, and the year they were introduced going from 1970 to 2020. The microprocessor data comes from this list. Intel 4004 is located in the lower left corner of the diagram and as mentioned it has 2,308 transistors. The AMD Epyc Rome processor in the upper right corner has 39.54 billion processors. Notice that the number of processors is on a logarithmic scale (not 1,2,3,4,5….but 1,10,100,1000,10000, etc.) A curve that grows exponentially becomes like a line in such a diagram. This is another way of saying that the technological progress for microprocessors is exponential. Every six and a half years the number of transistors is 10-doubled.

The diagram feature 100+ CPU processors. Time of introduction is along the x-axis and the number of transistors on the y-axis. | Transistors per Microchip Increased by More than 50 million Times in 50 Years
Moore’s Law: The number of transistors on microchips has doubled every two years. Moore’s law describes the empirical regularity that the number of transistors on integrated circuits doubles approximately every two years. This advancement is important for other aspects of technological progress in computing – such as processor speed or the price of computers. Data source: Wikipedia (wikipedia.org/wiki/Transistor_count). OurWorldData.org – Research and data to make progress against the world’s largest problem. Licensed under CC-BY by the authors Hannah Ritchie and Max Roser.

However, the processor that currently has the most transistors is the Cerebras Wafer Scale Engine 3 (WSE-3) with 4 trillion transistors. It is not in the diagram. A comparison between Intel 4004 and Cerebras Wafer Scale Engine 3 (WSE-3) may not be entirely fair since it is a wafer sized CPU specifically made for AI and much bigger than Intel 4004. However, the standard consumer or workstation CPU with currently the most transistors is Apple’s M3 Ultra (normal sized) with 184 billion transistors. However, 184 billion transistors that is more than 79 million times as many transistors as the Intel 4004.

A logarithmic scale from 1000 to 10 billion or 100 billion and over a time scale from 1971 to 2021. The graph showing the number of transistors is nearly linear.
Moore’s law: The number of transistors per microprocessor. Data source: Karl Rupp, Microprocessor Trend Data (2022) CC BY

This is a super fact because it is an important fact that is a surprise to you if you did not know about Moore’s law and mind-blowing even if you did.

Computational Capacity / Speed

One benefit of transistors and other electronic components becoming much smaller is increased speed. The graph below shows that the fastest supercomputers in 1993 had a processing speed of 124 Gigaflops and in 2025 a processing speed of 1.81 billion Gigaflops, which is 14.6 million times faster. One gigaflop equals one billion floating-point operations per second. Multiplying 3.1415926536 times 2.7182818284 = 8.5397342225 is an example of a floating point operation. Note that 1.81 billion Gigaflops is 1.81 quintillion floating-point operations per second.

The left axis in the diagram is logarithmic going from 100 gigaflops to well over a billion gigaflops. The x-axis corresponds to the years going from 1993 to 2025. | Transistors per Microchip Increased by More than 50 million Times in 50 Years
Computational capacity of the fastest supercomputers. Number of floating-point operations carried out per second by the fastest supercomputer in any given year. This is expressed in gigaflops, equivalent to one billion floating-point operations per second. Data source: Dongarra et al. (2025) OurWorldinData/technological-change | CC BY

Cost of Memory

Much smaller and faster and transistors and other kinds of electronic components also come with additional benefits such as an extreme reduction in the price of memory, including the extreme reduction in the price of random access memory, RAM. In the diagram below (taken from this page ) the blue line shows that the price for one Terabyte of memory (RAM) in 1957 was 3,79 quadrillion dollars and now in 2023 it is only 1,088 dollars. 3,79 quadrillion dollars is a lot of money.

You may object that and point out that 3,79 quadrillion dollars is even more than our national deficit. So how can that be possible? The answer is that one terabyte of memory did not exist in 1957. 3,79 quadrillion dollars per terabyte is the same as 3.79 million dollars per kilobyte, but those are the numbers that make sense for 1957. The gigantic computers back then only had a few kilobytes of memory. You just need to do the conversion. Anyway, 3,79 quadrillion dollars versus 1,088 dollars corresponds to a reduction in price by 3,48 trillion times. For disk memory the price per terabyte went down from 87.5 billion in 1956 to 11 dollars in 2023, a reduction of 7.95 billion times.

The blue graph corresponds to random access memory RAM and is most of the time the most expensive. The brown graph is disk memory. The red graph is flash memory, and the green line is solid state memory. All graphs show a huge reduction in price over time.
Historical price of computer memory and storage. This data is expressed in US dollars per terabyte (TB), adjusted for inflation. “Memory” refers to random access memory (RAM), “disk” to magnetic storage, “flash” to special memory used for rapid data access and rewriting, and “solid state to solid-state” drives (SSDs). Data source: John C. McCallum (2023); U.S. bureau of Labor Statistics (2026). Note: For each year, the time series shows the cheapest historical price recorded until that year. This data is expressed in constant 2020 US$. OurWorldinData.org/technological-change | CC BY

Increased Computational Power and Artificial Intelligence

My super fact 88 states that “Artificial Intelligence is Not New” It goes back to at least 1943 with the creation of the first artificial neural network model, the first trainable (able to learn) neural network in 1957, and the foundation of the field of “Artificial Intelligence Research” in 1956. In 1986 a landmark paper was published by David Rumelhart, Geoffrey Hinton, and Ronald Williams which introduced the Rumelhart backpropagation algorithm, which is used today by modern AI and Large Language Models, such as ChatGPT. Geoffrey Hinton, who is said to be the father of Artificial Intelligence, received the Nobel Prize in physics in 2024. David Rumelhart and Ronald Williams were both dead and could therefore not receive the Nobel Prize.

So why did it take so long for today’s commercial LLMs to appear? The answer is, for the most part, that the computational power and huge memory required did not exist until recently. Sure, training these large AI systems require enormous amounts of data that comes from outside of the processors (internet) but you need enormous amounts of memory to store this data and an enormous computational processing power to run the algorithms such as Rumelhart’s backpropagation algorithm. For example, ChatGPT 3.5 and 4.0 feature 96 versus 120 hidden layers of neurons with hundreds of billions and trillions of parameters/neurons in total. The enormous increase in computer processing power and memory is an essential aspect of scaling up AI. Our World in Data has an article about this here.

The diagram shows a deep learning neural network featuring hidden layers, a couple of input/output layers, and large computer. | Transistors per Microchip Increased by More than 50 million Times in 50 Years
The dots in the diagram are neurons. With permission from kwholley63

Conclusion

Moore’s law, a purely observational law, states that the number of transistors on a microchip doubles roughly every two years. This also means that the computational processing power of processors has greatly increased and the cost of memory has dropped by a lot. This in turn has made the existence of modern Large Language Models such as ChatGPT possible.

Transistors per Microchip Increased by More than 50 million Times in 50 Years. Processing speed has increased 14.6 million times in 32 years. The reduction in price per Tera Byte has gone down 3,48 trillion times in 67 years.

If Moore’s law holds, we can expect that in twenty years the number of transistors per Microchip will increase by more than 1,000 times, that processing speed will increase roughly 29,000 times, and that memory RAM will be more than 6,000 times cheaper. These are my calculations based on the graphs above.




To see the other Super Facts click here

Artificial Intelligence is Not New

Superfact 88: The history of artificial intelligence (AI) began in antiquity, with stories of artificial beings. The first artificial neural network model was created in 1943. The Turing test was created in 1950. The field of “Artificial Intelligence Research” was founded as an academic discipline in 1956. The first trainable (able to learn) neural network was demonstrated in 1957.

Since then, artificial intelligence has come a long way. Did you hear about the computer that defeated the reigning world champion in chess? A computer finally defeated the supreme human intellect in the world in an intellectual field. Is this the end of humanity? Oh, wait, that was in 1997.

White female AI robot using a microscope in the scientific laboratory. | Artificial Intelligence is Not New
Artificial intelligence and research concept. Shutterstock Asset id: 2314449325 by Stock-Asso

The various recent launches of large language models such as ChatGPT, Gemini, Claude, Llama, Deep Seek, etc., have impressed many people but also fooled many people into thinking that Artificial Intelligence is a new invention. It is not. Artificial Intelligence has been around for a long time, and its past is filled with many success stories as well as disappointments. Click here  to see a timeline for Artificial Intelligence stretching from antiquity to 2025. For additional sources click here, here, here, or here.

I consider this a super fact because it is true, kind of important, and based on my personal experience I believe that the long old history of Artificial Intelligence is a surprise to many.

My Personal Experience with Artificial Intelligence

In 1986, when I was in college in Sweden, I took a class in the LISP programming language. LISP was the first Artificial Intelligence programming language, and it was invented in 1958. In 1987, as a university level exchange student, I took a class called Artificial Intelligence at Case Western Reserve University. The book we used was Artificial Intelligence by Elaine Rich published in 1983. This book and the course were focused on decision trees and rule based algorithms and did not even mention neural networks.

That same year I also took a class called Pattern Recognition which introduced neural networks to me. In 1986 a landmark paper was published by David Rumelhart, Geoffrey Hinton, and Ronald Williams which introduced the Rumelhart backpropagation algorithm. Geoffrey Hinton received the Nobel Prize in physics in 2024. David Rumelhart and Ronald Williams were both dead and could therefore not receive the Nobel Prize. The Nobel Prize was also given to John J. Hopfield, another pioneer in neural networks. He invented the Hopfield network. You can read more about neural networks and the Nobel Prize in physics in 2024 here.

The Rumelhart backpropagation algorithm was a giant leap forward for neural networks and for Artificial Intelligence and it is the algorithm used by ChatGPT and the other large language models. Geoffrey Hinton is often interviewed in media and often presented as the father of Artificial Intelligence. He is not, but he is responsible for arguably the greatest leap forward in neural networks, as well as Artificial Intelligence.

In class we used the Rumelhart backpropagation algorithm to read images with text. It is one thing to type in a character on a keyboard and quite another to have a computer identify a character in an image. We trained our primitive neural networks to recognize images of letters using the Rumelhart backpropagation algorithm. We coded the backpropagation algorithm using the C programming language over perhaps 100 neurons/parameters and a few hundred synapses/weights (in AI). It worked pretty well. In comparison, ChatGPT 4 is estimated to have 1 trillion neurons/parameters. Our class was among the first in the world to try out this, at the time, new algorithm and at the time I did not realize the importance of it.

Later I did research and I worked in the field of Robotics where I implemented various Artificial Intelligence algorithms but not neural networks. I have a PhD in Applied Physics and Electrical Engineering with specialty in Robotics. At my next workplace Siemens I used decision tree algorithms, also Artificial Intelligence but not neural networks.

What is a Neural Network

Three blue circles connected to two red circles via lines assigned weights.
A simple old-style 1950’s Neural Network (my drawing)

The first neural networks created by Frank Rosenblatt in 1957 looked like the one above. You had input neurons and output neurons connected via weights that you adjusted using an algorithm. In the case above you have three inputs (2, 0, 3) and these inputs are multiplied by the weights to the outputs. 3 X 0.2 +0 + 2 X -0.25 = 0.1 and 3 X 0.4 + 0 + 2 X 0.1 = 1.4 and then each output node has a threshold function yielding outputs 0 and 1.

To train the network you create a set of inputs and the output that you want for each input. You pick some random weights and then you can calculate the total error you get, and you use the error to calculate a new set of weights. You do this over and over until you get the output you want for the different inputs. The amazing thing is that now the neural network will often also give you the desired output for an input that you have not used in the training. Unfortunately, these neural networks weren’t very good, and they sometimes could not even be trained.

As mentioned, in 1986, Geoffrey Hinton, David Rumelhart and Ronald J. Williams presented the Rumelhart backward propagation algorithm which were applied to a neural network featuring a hidden layer (at least one hidden layer). It was effective and it was guaranteed to learn patterns that were possible to learn. It set off a revolution in Neural Networks. In the network below you also use the errors in a similar fashion as in the Rosenblatt network. However, the combination of a hidden layer and the backpropagation algorithm make a huge difference.

Three blue circles connected to four yellow circles connected to two red circles all via lines assigned weights.
A multiple layer neural network with one hidden layer. This set-up and the associated backpropagation algorithm set off the neural network revolution. My drawing.

Below I am showing two 10 X 10 pixel images containing the letter F. The neural network I created in class (see above) had 100 inputs, one for each pixel, a hidden layer and then output neurons corresponding to each letter I wanted to read. I think I used about 10 or 20 versions of each letter during training, by which I mean running the algorithm to adjust the weights to minimize the error until it is almost gone. Now if I used an image with a letter that I had never used before, the neural network typically got it right even though the image was new.

The 10 X 10 pixel images are filled with black pixels resembling two differently looking characters F | Artificial Intelligence is Not New
Two examples of the letter F in a 10 X 10 image. You can use these images (100 input neurons) to train a neural network to recognize the letters F.

At first, it was believed that adding more than one hidden layer did not add much. That was until it was discovered that by applying the backpropagation algorithm differently to different layers created a better / smarter neural network and so at the beginning of this century the deep learning neural networks were born (or just deep learning AI). I can add that our Nobel Prize winner Geoffrey J. Hinton was also a pioneer in deep learning neural networks.

Three blue circles connected to four yellow circles connected to four green circles connected to six blue circles connected to two red circles all via lines representing weights.
My drawing of a deep learning neural network (deep learning AI). There are three hidden layers.

I should mention that there are many styles of neural networks, not just the ones I’ve shown here. Below is a network called a Hopfield network (it was certainly not the only thing he discovered).

Four neurons that are all connected to each other.
In a Hopfield network all neurons are input, and output neurons and they are all connected to each other.

For your information, ChatGPT-3.5 and ChatGPT-4 are deep learning neural networks, like the one in my colorful picture above, but instead of 3 hidden layers it has 96 hidden layers in its neural network and instead of 19 neurons it has a total of 176 billion neurons.

The Dark Side of AI

The potential harm of AI is a related and important topic that I did not address. However, this is already a very long and complex post, and I don’t know enough about this topic (yet). To read more about this topic check the comments made by “Grant at Tame Your Book” (in comment section). Better, Grant wrote and excellent, well research and professional post about this issue called Don’t Confuse AI with a Benign Tool. Please check it out.




To see the Other Super Facts click here