mej7.com How AI Works

mej7.com — laborator falas interaktiv për të mësuar si funksionon AI

mej7.com është një platformë edukative falas që shpjegon, me vizualizime, si kalon teksti nga fjalët e tua te tokenët, embedding-et, attention, transformer-i dhe gjenerimi i përgjigjes. Është për fillestarë, nxënës dhe kureshtarë teknikë — pa pagesë për përmbajtjen në browser. Përmbajtja ofrohet në shqip dhe anglisht (mej7.com/en/).

Rreth projektit

Çfarë është: një faqe e vetme me laborator interaktiv dhe tre nivele — Beginner, Intermediate (tokenë, embeddings, Neural Network Lab, attention, training, quiz) dhe Advanced (backprop, KV cache, MoE, RLHF, kuantizim).

Për kë: njerëz që duan të kuptojnë si funksionon AI / modelet gjuhësore, veçanërisht në shqip.

Çfarë e dallon: vizualizime interaktive hap-pas-hapi; kurrikula e strukturuar në faqe (shih content/roadmap.json në repozitor).

Pyetje të shpeshta

A duhet të paguaj?

Jo. Përmbajtja edukative në faqe është falas. Regjistrimi opsional shërben për ruajtjen e progresit (API në repozitor).

A më duhet programim?

Jo për fillestarët. Niveli Advanced supozon kuriozitet për matematikë dhe kod.

A është faqe zyrtare e OpenAI?

Jo. Projekt edukativ i pavarur; vizualizimet janë ilustruese, jo peshë reale nga modele prodhimi.

mej7.com
A më duhet Python ose programim?

Jo për fillestarët. Laboratori Intermediate punon në browser. Niveli Advanced supozon kuriozitet për matematikë dhe kod.

A është kjo faqe e një kompanie AI?

Jo. Është projekt edukativ i pavarur. Vizualizimet janë ilustruese; disclaimer-et në faqe theksojnë që numrat nuk janë nga modele prodhimi.

Si ndryshon gjuha?

Përdor flamujt 🇬🇧 / 🇽🇰 në header për anglisht dhe shqip. URL opsionale: https://mej7.com/?lang=en.

About (English)

mej7.com is a free bilingual (Albanian & English) interactive site for learning how language models work — tokens, embeddings, attention, transformers, training, plus an advanced track. Optional login syncs progress; the lab works without an account.

FAQ (English)

Cost? Free educational content in the browser.

Official OpenAI product? No — independent teaching site with simplified visuals.

A VISUAL AI LABORATORY

What actually happens inside AI?

From your words to tokens, vectors, matrices, attention, probabilities and the final answer — explore what happens inside a modern language model.

// Don't just use AI. Understand what happens inside it.
FOR EVERYONE — EVEN A 4-YEAR-OLD

How AI works, in the simplest way possible

Forget the hard words for a second. AI does three simple things, over and over, super fast: it Looks, it Guesses, and it Says. Watch the boxes spin.

TOKENIZATION

Your words become tokens

Neural networks don't process words as human concepts. Text is first converted into tokens — pieces of text represented numerically.

Text↓Tokens↓Token IDs

Token IDs depend entirely on a model's specific tokenizer/vocabulary — the IDs shown here are illustrative, not universal.

SUBWORD TOKENIZATION

Long or rare words are often split into smaller pieces. For example, "unbelievable" might become:

unbelievable

This is just an illustrative example — real splitting depends on the specific tokenizer in use.

EMBEDDINGS

Tokens become vectors

Each token ID is looked up in a table and turned into a vector — a list of numbers. Words with related meanings tend to end up with nearby vectors.

Token ID→Embedding lookup→Vector
ILLUSTRATIVE VALUES
"king"[0.21, -0.42, 0.87, 0.13, ...]

The embedding space below is reduced to 2 dimensions for human eyes — real models use hundreds or thousands of dimensions.

Simplified visualization for teaching purposes — positions are illustrative, not real coordinates from a production model.

VECTOR ARITHMETIC

This is the classic illustration of what embeddings capture: relationships between words behave a bit like arrows you can add and subtract.

king−man+woman≈queen
king man woman result ≈ queen

Illustrative geometry in a simplified 2D space — real embedding arithmetic happens across hundreds of dimensions and doesn't always land this cleanly.

WEIGHTS

The model is full of numbers called weights

Every connection between neurons has a weight. Change the weights below and watch the output change live. Click "Run forward pass" to see values flow through the network.

CONTRIBUTION
Input = 0.80×Weight = 0.50=0.40

This is the basic formula: input × weight = contribution. Contributions from many inputs are summed to form a neuron's activation.

· NEURAL NETWORK LAB

See the network compute

Full interactive feed-forward laboratory — forward pass, matrix multiply, training on real toy data, and neuron autopsy. All numbers come from the actual simulation.

MATRICES

Neural networks multiply numbers

Matrix multiplication lets a network transform large amounts of information in a single mathematical step. Click a result cell to see exactly which row and column combined to produce it.

MATRIX A
×
MATRIX B
=
RESULT C

Amber = the row from A being used. Cyan = the column from B being used. Each result cell is the dot product of that row and that column.

· LINEAR ALGEBRA

What a matrix really does: it bends space

Every 2×2 matrix is a rule for transforming every point in a plane — stretching, rotating, shearing, or flipping it. Drag the sliders, or try a preset, and watch the whole grid warp in real time.

Determinant (area scale factor)1.00

Positive determinant: space keeps its orientation — nothing gets flipped.

î (where "right" ends up) ĵ (where "up" ends up)
ACTIVATION

Activation functions

After each multiplication, the result passes through an activation function, which lets the network learn non-linear patterns.

f(x)0.00
ATTENTION

What should the model focus on?

Pick a word and see how much "attention" it pays to every other word in the sentence — both as connecting lines and as a heatmap.

Attention doesn't simply mean "the AI understands what's important" — it's a mathematical mechanism that lets every token incorporate information from other tokens, with varying strengths.

Q / K / V

Query, Key, Value

Every token produces three vectors with different roles in the attention mechanism. The angle between a Query and a Key vector determines how strongly they match.

Q
Query
What this token is looking for
K
Key
What each token offers for matching
V
Value
Information passed forward
Similarity (dot product)—
Q × Kᵀ↓attention scores↓softmax↓weighted combination of V
TRANSFORMER

The transformer block

Click each block to see its explanation. Many blocks like this are stacked on top of each other.

TRANSFORMERS — Q·K·V IN DETAIL

Inside the Transformer: how Q, K and V actually multiply

A full, real walk-through of scaled dot-product attention: computing Q, K and V, transposing K, multiplying, scaling, applying the triangular causal mask, and running softmax — with real numbers, for a tiny 4-token example.

X = 4 tokens × 2 dims Wq, Wk, Wv = 2 × 2
STEP1 / 12

Numbers here are a small, made-up example chosen to be easy to follow by hand — real models use hundreds of dimensions and different values, but the operations (multiply, transpose, scale, mask, softmax) are exactly these.

This triangular mask is used in autoregressive (decoder-style) models like GPT, which generate text left-to-right. Encoder-style models (like BERT) typically skip this mask and let every token see the whole sequence.

MULTI-HEAD ATTENTION

Several attention "heads" at once

The model runs several attention mechanisms in parallel ("heads"), each potentially learning different patterns.

Different heads can learn different patterns — the examples shown here are simplified interpretations for visualization, not fixed roles proven for every model.

LOGITS & SOFTMAX

How does AI choose the next word?

The model produces a number ("logit") for every possible token, then softmax turns them into probabilities.

Low temperature → more concentrated distribution. High temperature → more spread out / random. This doesn't make the AI "smarter" or "dumber" — it just changes the randomness of the choice.

AUTOREGRESSIVE GENERATION

Next-token generation simulator

The model generates one token at a time, adds it to the context, and repeats. Click a candidate token to add it.

CURRENT CONTEXT
Context↓Tokenization↓Model↓Logits↓Softmax↓Choose Token↓Add To Context↻
TRAINING

But where did all these weights come from?

Weights are learned by comparing the model's predictions with the real target and gradually adjusting.

Massive dataset↓Prediction↓Compare with target↓Loss↓Backpropagation↓Update weights↻
SIMPLIFIED EXAMPLE
Target token"cat"
Loss (simplified)—
weight0.7200
gradient—
new weight—
GRADIENT DESCENT

The loss landscape

Training moves the "ball" toward lower-loss regions, one step at a time.

PARAMETERS

What does 7B or 70B parameters mean?

A parameter ≈ one learned numerical value (a weight). More parameters means more numbers to fit — but that alone doesn't guarantee higher "intelligence".

Number of parameters1,000,000

Parameter count doesn't directly equate to intelligence — architecture, data, training method and optimization all matter too.

CONTEXT WINDOW

How much the model can "see" at once

The model can only see a limited number of tokens at a time — the context window.

Context windows are measured in tokens and vary by model.

HALLUCINATIONS

Why can AI be confident and wrong?

The model generates statistically likely sequences, rather than querying a database of guaranteed facts.

QUESTION: "The capital city of the planet Mars is..."

The model can produce a confident-sounding answer even when no correct answer exists — because of limits in training data, ambiguous prompts, statistical generation, and the lack of guaranteed factual verification. This isn't simply "the AI doesn't know."

PROMPT ENGINEERING

Write prompts that actually work

Models respond to instructions + context + format. Compare a vague prompt with a structured one — same task, very different results.

❌ Vague prompt

✓ Structured prompt

Scores are a teaching rubric (clarity, context, output format) — not output from a live model.

FOLLOW ONE TOKEN

Watch a single token travel through the whole model

Pick a word below, then step through exactly what happens to it — from raw text to its contribution to the final answer.

STEP1 / 9

This journey uses simplified, illustrative numbers to make the trip easy to follow — the sequence of operations (tokenize, embed, attend, transform, predict) mirrors real models.

EXCLUSIVE — LIVE VISUALIZATION

The Living Network — watch AI think, live, as you type

A 68-neuron network that reacts to every single keystroke in real time. Every letter you type sends a pulse of light into the network — watch it ripple, cascade and fade across the layers, exactly like the signal-passing you learned about above.

LIVE · REACTS TO YOUR TYPING
NETWORK ENERGY
0%

This is a real, tiny feed-forward simulation (fixed random weights, sigmoid-like squashing, decay over time) driven live by your keystrokes — built purely to be watched. It is not a language model and doesn't "understand" what you type; it's a window into the kind of signal-passing every real neural network does, made visible.

MECHANISTIC INTERPRETABILITY

Look inside: which words does the model attend to?

Real transformers use "attention" to decide, for every word, how much to look at every other word. This is a simplified but structurally faithful attention map — type a sentence and see which tokens light up for each other, the same kind of picture mechanistic-interpretability researchers stare at all day.

SELF-ATTENTION · TOY MODEL

Rows are "query" tokens, columns are "key" tokens; brighter cells mean stronger attention weight (softmax over a toy scoring function keyed on token content and distance). Real attention heads are learned from data and often specialize — one head might track subjects and verbs, another might track punctuation. This toy version is for intuition, not a real model's internals.

3D EMBEDDING SPACE

Words live as points in space — rotate and see

Every token becomes a vector of numbers — a point in a high-dimensional space where similar meanings sit close together. Here's a 3D slice of that idea, spinning slowly.

Positions here are illustrative, not computed from a real embedding model — but the core idea is exactly this: related words (animals, actions, places) cluster together in the space a transformer actually reasons in.

FLASHCARDS

Drill the core vocabulary

Flip a card, test yourself, and move on. Your streak of known cards is saved in this browser.

Known this session0
KNOWLEDGE CHECK

Do you understand it now?

A short 6-question check covering everything above. Your best score is saved in this browser.

Question1 / 6
Best score—

· VISUALIZATION ENGINE

One engine, many labs

Probability bars, attention heatmaps, pipeline flows, and tensor shapes — shared building blocks behind the interactive lessons you have already explored.

Softmax → probability bars

Drag temperature: sharp peaks mean the model is confident; flat bars mean many tokens are plausible.

Attention matrix

Each row is one token “looking at” the others. Brighter cells = stronger connection (toy demo, same math as the mechanistic lab).

Tensor shapes

When docs say [batch, seq, dim], they mean stacked lists of numbers — not magic boxes.

Mini pipeline

The full recap at the end of the lab uses the same flow renderer — click a stage.

· COMPLETE PIPELINE

The whole journey, start to finish

You've now seen every piece up close. Here's the whole thing laid out in one line — click any step to recall what it does.

DAILY CHALLENGE

One quick question. Every day.

Build a micro-habit — answer today's AI question and keep your challenge streak alive.

Challenge streak1

ACHIEVEMENTS

Collect badges as you master AI

Explore sections, keep streaks, ace the quiz, and sync your account to unlock them all.

AI TIME MACHINE

Scrub through AI history

From the first neuron to ChatGPT — see the milestones that led to today's language models.

2024

LEARNING MAP

Your constellation of progress

Each dot is a section in this lab. Orange stars mean you've been there — light up the whole sky.

You've seen what happens inside the model. Now try changing it.

© 2026 How AI Works — An educational page. All visualizations are simplified for teaching purposes and do not represent real weights or activations from any production model.