Skip to content

Instantly share code, notes, and snippets.

@ashokvarmamatta
Last active June 1, 2026 05:51
Show Gist options
  • Select an option

  • Save ashokvarmamatta/004d82fe524d1bf514df4f23cd0cbe17 to your computer and use it in GitHub Desktop.

Select an option

Save ashokvarmamatta/004d82fe524d1bf514df4f23cd0cbe17 to your computer and use it in GitHub Desktop.
Neurons Explained From Scratch - From a single weighted decision to production on-device AI (Python + PyTorch + Gemma). For ECE/CS grads new to ML.

๐Ÿง  Neurons Explained From Scratch

From First Principles โ†’ Production On-Device AI

Typing SVG

Karpathy Python PyTorch On--Device AI


๐ŸŒฑ Heads up โ€” I'm still learning. I'm an ECE grad new to ML, working through Karpathy's "Zero to Hero" from scratch and writing as I go. Things here may be incomplete or slightly off. If you spot a mistake, please open an issue / drop a comment โ€” I'd rather get it right than be polite about being wrong. The point is the journey, not authority.

๐Ÿ“š Chapter 1 of an ml-from-zero series. Each chapter follows the same pattern: intuition โ†’ math โ†’ Python โ†’ PyTorch โ†’ on-device relevance. The full series repo goes public once a few more chapters land. Star this gist if it helped, and follow my GitHub profile to see new chapters as they ship.


๐Ÿ“– Who this is for

You finished engineering (any branch โ€” I did ECE), never touched ML, and people keep talking about neurons / weights / backprop like it's obvious. This guide goes from "a complete beginner can follow it" to "you can ship one on a phone" without skipping steps.

What you'll know by the end:
   โœ“ What a neuron actually computes (with real numbers)
   โœ“ Why we "squish" (and why neural nets need it)
   โœ“ How to write a neuron in Python from scratch โ€” 6 lines
   โœ“ The same neuron in PyTorch โ€” 1 line
   โœ“ Why this matters for on-device AI (Gemma, AICore, NPU)

No prior ML. No calculus required (yet). Just one decision, one formula, one Python file.


๐Ÿฆ Step 1 โ€” The intuition (no code, no math)

Forget code. Forget math. Just a person deciding whether to eat ice cream right now.

SHOULD I EAT ICE CREAM RIGHT NOW?

   Am I hungry?       โ†’  4 out of 10
   Is it hot outside? โ†’  8 out of 10
   Do I have money?   โ†’  5 out of 10

But these inputs don't matter equally. Hunger matters a LOT. Heat matters some. Money barely matters (it's just ice cream).

So you give each input a weight โ€” a number for "how much do I care":

   Hunger  ร— 3   (matters a lot)
   Heat    ร— 2   (matters some)
   Money   ร— 1   (barely)

Then multiply and add them up:

   (4 ร— 3) + (8 ร— 2) + (5 ร— 1) = 12 + 16 + 5 = 33

Plus a small bias โ€” your default mood toward ice cream today (+2 because it's been a good day):

   33 + 2 = 35

35 is high โ†’ YES, eat ice cream.

That is a neuron. Inputs โ†’ multiply by weights โ†’ add them up โ†’ add a bias โ†’ get a number โ†’ use that number to decide.


๐Ÿงฎ Step 2 โ€” Write it down (the formula)

In math notation:

   y = wโ‚ยทxโ‚ + wโ‚‚ยทxโ‚‚ + wโ‚ƒยทxโ‚ƒ + b

Where:

  • xโ‚, xโ‚‚, xโ‚ƒ = the inputs (hunger=4, heat=8, money=5)
  • wโ‚, wโ‚‚, wโ‚ƒ = the weights (3, 2, 1 โ€” how much each input matters)
  • b = the bias (+2, the default mood)
  • y = the output (35 in our example)

That's it. That's the whole neuron โ€” multiplied, added, biased.

"Where do the weights 3, 2, 1 come from?" Right now we're hand-picking them. In a real network the weights and bias start random and get nudged toward useful values automatically โ€” that's training, driven by backpropagation. Both covered in the next chapter; for now, just take the numbers as given.


๐ŸŒ€ Step 3 โ€” The "squish" (this is the key thing)

35 is not a decision. You need YES or NO, not a number.

So neurons add one more step: squash the result into a small range (usually -1 to +1, or 0 to 1). This is called an activation function.

The most common ones:

Name What it does Output range
tanh (S-curve) Smoothly bends any number to fit between -1 and +1 -1 to +1
ReLU Negative โ†’ 0, positive โ†’ unchanged 0 to โˆž
sigmoid Smoothly bends any number to fit between 0 and 1 0 to 1

For our example:

   tanh(35) โ‰ˆ 1.0    โ†’  +1   โ†’  "YES eat ice cream"
   tanh(-5) โ‰ˆ -1.0   โ†’  -1   โ†’  "NO don't eat"

(tanh saturates fast โ€” by ยฑ2 you're already at ~ยฑ0.964, by ยฑ3 you're essentially at ยฑ1.)

Why bother squishing?

This is the part most beginners skip and regret later. Without the squish, stacking 100 neurons gives you the same power as 1 neuron.

Quick proof. Without squish, each neuron is just multiply + add. Stack three:

   Neuron 1:  y = 2x + 1
   Neuron 2:  y = 3ยท(prev) + 5  โ†’  3(2x+1) + 5 = 6x + 8
   Neuron 3:  y = 4ยท(prev) โˆ’ 2  โ†’  4(6x+8) โˆ’ 2 = 24x + 30

Three layers collapsed into one (24x + 30). Stacking did nothing.

With the squish (tanh), the algebra refuses to simplify. Each layer keeps its own work. That's why deep networks need activation functions. The fancy word: nonlinearity.

   Without nonlinearity:  100-layer net = 1-layer net
   With nonlinearity:     each layer earns its keep

๐Ÿ’ป Step 4 โ€” Write a neuron in pure Python (6 lines)

No libraries. Just math.

import math

def neuron(inputs, weights, bias):
    # 1. Weighted sum
    s = sum(w * x for w, x in zip(weights, inputs)) + bias
    # 2. Squish through tanh
    return math.tanh(s)

# Run the ice-cream decision:
inputs  = [4, 8, 5]      # hunger, heat, money
weights = [3, 2, 1]      # how much each matters
bias    = 2              # default mood

y = neuron(inputs, weights, bias)
print(y)                 # 1.0   โ†’  YES, eat ice cream

โœ… Verified: ran this on Python 3.10 (stdlib only, no installs) โ€” output is 1.0 as shown.

That's a real working neuron. You just built one.

Try changing inputs = [1, 0, 0] (barely hungry, not hot, no money). You still get y โ‰ˆ 0.9999 โ€” the weighted sum is 1ยท3 + 0ยท2 + 0ยท1 + 2 = 5, and tanh(5) โ‰ˆ 0.9999. Even weak inputs combined with the bias push the sum past the point where tanh saturates to +1. That's the bias doing its job: it nudges the decision when inputs alone are not strong enough.


๐Ÿ”ฅ Step 5 โ€” The same neuron in PyTorch (1 line)

PyTorch is the most common ML framework in research and many industry teams. Alternatives exist โ€” Google and Anthropic mostly use JAX, and on-device often uses TensorFlow Lite or ONNX Runtime โ€” but PyTorch is the safest "where do I start?" answer. The neuron we just built becomes:

import torch

# A "Linear" layer with 3 inputs, 1 output = exactly one neuron
neuron = torch.nn.Linear(in_features=3, out_features=1)

# Initialise weights and bias (PyTorch does this randomly by default)
with torch.no_grad():
    neuron.weight.copy_(torch.tensor([[3.0, 2.0, 1.0]]))
    neuron.bias.copy_(torch.tensor([2.0]))

# Run the same ice-cream decision
inputs = torch.tensor([[4.0, 8.0, 5.0]])
output = torch.tanh(neuron(inputs))
print(output.item())     # 1.0

โœ… Verified: installed PyTorch 2.12 (CPU) and ran this โ€” it prints 1.0, exactly matching the pure-Python neuron in Step 4. Same neuron, bigger toolbox.

Same answer. Same neuron. PyTorch just gives you:

  • GPU acceleration (this neuron runs on GPU automatically)
  • Vectorisation (process 1000 ice-cream decisions at once)
  • Autograd (it remembers HOW you computed y so it can adjust weights via backprop later โ€” but that's a different post)

The math is identical. The toolbox is bigger.


๐Ÿ—๏ธ Step 6 โ€” Why one neuron isn't enough

One neuron can decide one thing based on weighted inputs. That's powerful but limited. Real problems need many decisions stitched together:

   IMAGE OF CAT
        โ†“
   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
   โ”‚ Layer 1: 64 neurons                 โ”‚  โ† roughly: edges / corners
   โ”‚ Layer 2: 32 neurons                 โ”‚  โ† roughly: shapes (eye, ear, fur)
   โ”‚ Layer 3: 16 neurons                 โ”‚  โ† roughly: parts (head, tail, paw)
   โ”‚ Layer 4: 2 neurons โ†’ [cat, dog]     โ”‚  โ† final classification
   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
        โ†“
   "cat with 85% confidence"

That stack is called a neural network. Each layer is a row of neurons all reading the same input. Stack more layers = "deep learning."

๐Ÿ’ก Quick side-note: the "edges โ†’ shapes โ†’ parts" hierarchy is really the convolutional net intuition โ€” produced by the layer wiring (filters that slide over the image), not by a single neuron. Dense neurons on tabular data don't form a clean visual hierarchy like this. The underlying neuron is the same; the wiring on top differs.

But every single neuron in there is the exact thing we just built:

   y = activation(wโ‚ยทxโ‚ + wโ‚‚ยทxโ‚‚ + ... + b)

GPT-5, Gemma, every image classifier, every speech recogniser โ€” they all use this neuron as their atomic unit inside dense / feed-forward (FFN) layers. What changes between them is the wiring on top (attention, convolutions, normalisation, residual connections) and the scale (billions of these neurons instead of dozens).

โš ๏ธ One caveat for transformers: the attention mechanism in GPT-5 / Gemma / Claude is NOT a stack of these neurons โ€” it's softmax(QยทKแต€ / โˆšd) ยท V, a different matrix-multiply + softmax operation. The FFN sub-layers between attention blocks ARE made of these neurons; attention itself isn't. So "every modern AI is built on this neuron" holds for image classifiers and the FFN parts of transformers, but attention deserves its own chapter.


๐Ÿ“ฑ Step 7 โ€” Why this matters for on-device AI

A growing share of AI work is moving to the phone, not just to cloud datacentres. Google ships Gemma models that run on Pixel via AICore. Apple runs models on the Neural Engine. Qualcomm Snapdragon NPUs run small LLMs on-device too (exact tokens/sec depends on model size, quantization, and chip generation โ€” check the device's own benchmarks).

For on-device models, every parameter costs battery + memory. Quick (rough) back-of-envelope:

  • A small on-device LLM today has on the order of ~2 billion parameters (the weights and biases learned during training โ€” note: parameters โ‰  neurons; each neuron contributes several parameters).
  • At 4 bytes per parameter (float32), that's 2 ร— 10โน ร— 4 = 8 GB of RAM โ€” too much for most phones.
  • Quantization shrinks each weight to ~4 bits (~0.5 bytes) โ†’ roughly 2 ร— 10โน ร— 0.5 = 1 GB. Now it can fit on a flagship phone.
  • The math gets noisier (less precise weights) but the network typically still works because most weights only need rough values.

That's the gap between "AI works" and "AI works on your phone in your pocket without burning the battery." Engineers who understand the neuron at this level can:

  • Profile NPU vs GPU performance per layer
  • Reason about quantization error per neuron
  • Pick activations that play well with INT8 hardware

This is the on-device AI engineer niche. It starts with knowing what a neuron is.


๐ŸŽฏ Recap โ€” the whole post in 30 seconds

Concept One-liner
Neuron One weighted decision: tanh(wโ‚xโ‚ + wโ‚‚xโ‚‚ + ... + b)
Weight How much each input matters (set by training)
Bias Baseline shift โ€” moves the decision threshold even when inputs are weak
Activation (tanh/ReLU) The squish that bends the output; lets stacking layers actually add power
Nonlinearity The fancy word for "the squish prevents collapse"
Layer A row of neurons all reading the same input
Network A stack of layers (Gemma, GPT, every image classifier)

That's the atomic unit. Everything else in modern AI โ€” transformers, CNNs, diffusion models, on-device LLMs โ€” is wiring, scale, and training built on top of this one weighted decision.


๐Ÿ”— Next steps

  1. Type the Python neuron above yourself โ€” actually run it. Change the bias from +2 to -10 and see the decision flip.
  2. Read Karpathy's "Zero to Hero" Lecture 1 โ€” https://www.youtube.com/watch?v=VMj-3S1tku0. Builds micrograd (a baby PyTorch) from this same neuron, end-to-end in 2.5 hours.
  3. Drill these 28 questions cold โ€” coming in the next post.
  4. Build the Jupyter notebook walkthrough โ€” that's where I am next.

๐Ÿ“˜ Glossary โ€” jargon used in this post (click to expand)
  • NPU (Neural Processing Unit) โ€” A dedicated chip block inside modern phone SoCs (Pixel Tensor, Apple A-series, Qualcomm Snapdragon) for running neural-network math. Optimized for matrix multiply + activation functions at low power; faster and more energy-efficient than CPU/GPU for inference.

  • INT8 โ€” 8-bit integer representation (-128 to +127). Compared to float32 (32 bits per number), it makes models ~4ร— smaller and faster on hardware that supports INT8 arithmetic. Less precise per number, but networks usually still work because most weights only need rough values.

  • AICore โ€” Google's on-device AI runtime on Pixel devices. Apps query it like a system service; under the hood it dispatches to NPU/GPU/CPU based on the model and hardware available.

  • Neural Engine โ€” Apple's NPU, embedded in every iPhone/iPad SoC since A11 (2017). Apple's equivalent of the NPU block in Google Tensor or Qualcomm Snapdragon.

  • Quantization โ€” Converting model weights from float32 (4 bytes per number) to INT8 (1 byte) or 4-bit (~0.5 bytes). Memory shrinks 4โ€“8ร—, runs faster on NPUs, costs a small accuracy hit. The reason a multi-billion-parameter model can fit on a phone.

  • float32 / FP32 โ€” 32-bit floating-point representation, the default precision for weights during training. 4 bytes per number. "float16" / "bfloat16" / "int8" / "int4" are smaller alternatives used for inference.

  • Autograd โ€” Short for "automatic differentiation." The library code that auto-computes gradients for backprop, so you don't derive them by hand. PyTorch's torch.autograd, JAX's grad, TensorFlow's GradientTape all do this. Karpathy's micrograd is a ~100-line educational autograd you can read end-to-end.

  • FFN (Feed-Forward Network) โ€” In a transformer block, the dense sub-layer that sits between attention layers. Made of the same neurons we built here, just lots of them in parallel. Two linear layers with a non-linearity between them.

  • Attention โ€” The core transformer operation: softmax(QยทKแต€ / โˆšd) ยท V. Lets each token weigh how much to "look at" every other token. Distinct from a neuron โ€” it's a matrix multiply + softmax, not a weighted sum + tanh. Covered in its own future chapter.


Built from notes while studying Karpathy L1 as an ECE-grad new to ML.
๐Ÿ‘‰ Chapter 2 is live: Backprop & the Chain Rule From Scratch โ€” how the weights actually learn, with a 40-line micrograd engine verified against PyTorch.


๐Ÿ’ฌ Found a mistake? Help me learn. If a sentence is wrong, a number's off, or a framing is misleading โ€” please open an issue or drop a comment. Same if a concept didn't click; tell me where it broke and I'll rewrite that section.


Author: Matta Ashok Varma ยท ECE grad โ†’ Android dev (4+ yrs ยท 18 production apps) โ†’ on-device AI ยท learning in public.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment