Skip to content

API Reference

Auto-generated from the docstrings of the public optimumai package.

Top-level exports

Everything in optimumai.__all__ is the public API and follows semantic versioning:

# Linear algebra
from optimumai import Vector, Matrix

# Probability & calculus
from optimumai import softmax, derivative, gradient, integrate

# Autograd
from optimumai import Value

# Neural networks
from optimumai import MLP, Adam, minimize

# Transformers
from optimumai import Attention, MultiHeadAttention, TransformerBlock, TextPipeline

# World models & interpretability
from optimumai import JEPA, superposition

# Embeddings, RAG, diffusion
from optimumai import embedding_lookup, nearest_neighbors, RAGPipeline, forward_diffusion

# Systems
from optimumai import kv_cache_size, vram_estimate

# Generation & tutor
from optimumai import generate

# Course & progress
from optimumai import COURSE, ProgressTracker, Quiz, ReviewScheduler

# Visualization
from optimumai import render_concept, editable_plot

# Kernels
from optimumai import KernelWorkbench, GpuSim

# Trace types
from optimumai import Trace, ExplainLevel, Lesson, Course, Workbook

Module reference

optimumai

OptimumAI — unlock the math behind AI.

Every operation, from a dot product to a full transformer block, can be run with explain=True to produce a step-by-step computation trace, a terminal visualization, and the context for why AI uses it.

>>> from optimumai import Vector
>>> Vector([1, 2, 3]).dot(Vector([4, 5, 6]), explain=True)
32.0

v0.2 adds the fundamentals behind modern AI: a micrograd-style autograd engine, calculus, optimizers, neural networks with real backprop, multi-head attention, transformer blocks, LeCun's JEPA world model, and Anthropic-style superposition.

v0.3 turns it into a learning path: a structured Course, progress tracking, a Streamlit dashboard, plus embeddings, RAG, diffusion, and an optional LLM tutor.

v0.4 adds the foundations of the stack: tensors & integration, the PyTorch and JAX programming models, and the systems layer — the CUDA execution & memory model, tiled matmul kernels, the KV cache, and a VRAM budget calculator.

v0.5 makes it interactive: feed your own input and watch it flow — a REPL, a text-to-transformer pipeline, op comparisons, parameter sweeps, and symbolic differentiation of your own equations.

v0.6 adds graphs: matplotlib figures (via optimumai.visualization) for activation curves, attention heatmaps, embedding scatter, and a 3D loss landscape with the gradient-descent trajectory carved across it.

v0.7 adds the circuit: render any expression or Value graph as a computation "circuit" (interactive HTML, Graphviz DOT, or terminal) via optimumai.circuit, with data and gradients flowing the wires.

v0.8 adds frontier concepts: FlashAttention (IO-aware tiling + online softmax), quantization (int8/int4), LoRA (parameter-efficient fine-tuning), and DPO (preference alignment) — how today's large models are actually built and run.

v0.9 makes it a learning product, grounded in cognitive science: a quiz / active-recall engine (the testing effect), spaced-repetition review (SM-2), guided onboarding, and course search.

v0.10 is hands-on: GPU kernels from scratch on a pure-Python simulator (write & grade your own), an editable equation↔graph in the browser, animated GIF export, and compute-the-answer exercises. Get your hands dirty.

v1.0 is the stable release: real token generation (Ollama / Hugging Face / Anthropic / a toy fallback), interactive drag-the-inputs circuits, a visualize-any-concept registry (PNG + GIF), a notebook launcher, and a docs site.

v1.1 broadens from the deep-learning stack to the whole field: classical machine learning (regression, trees, k-means, KNN, naive Bayes, PCA, metrics), classical AI search (BFS/DFS/UCS, A*, minimax + alpha-beta), reinforcement learning (MDPs & the Bellman equation, Q-learning/SARSA, REINFORCE, PPO), NLP fundamentals (BPE, TF-IDF, n-grams, edit distance, word2vec), computer vision (convolution, pooling, Sobel edges, a tiny CNN), and LLM evaluation (BLEU/ROUGE, perplexity, calibration, and a candid hallucination/faithfulness heuristic) — each an explainable trace.

v1.2 makes it interactive & explained: prompt-engineering patterns (zero/few-shot, chain-of-thought, ReAct, self-consistency, structured output), the augmented-RNN lineage from distill.pub (attention as differentiable memory, Neural Turing Machines, Adaptive Computation Time), self-contained interactive HTML playgrounds (a Transformer-Explainer-style attention widget, plus k-means and A*), and a per-concept plot/GIF gallery wired into the visualize registry.

v1.3 makes it flow: a Plot Studio (feed numbers, get any chart — bar, histogram, scatter, box, line, pie, violin — plus the exact matplotlib + numpy code on screen), and a flows subpackage of distill.pub-style interactive circuit-flow diagrams (a transformer forward pass, scaled dot-product attention, TF-IDF, and word2vec) rendered as self-contained, offline HTML.

v1.4 teaches the tools: runnable, explained tutorials for NumPy, matplotlib, PyTorch (built on OptimumAI's own engine so it runs torch-free, with the real torch code shown), and LLM fine-tuning (a numpy SFT → LoRA → DPO toy pipeline plus the production HF/PEFT/TRL code). Walk any of them in the terminal (optimumai tutorial numpy) or export to a notebook — and read the matching in-depth guides on the docs site.

v1.5 hardens the interactive layer with OptiX, a typed, unit-tested TypeScript widget kit (authored in web/, compiled by esbuild to a single self-contained IIFE shipped in the wheel) — replacing hand-written inline JS. Its debut is a TensorFlow-Playground-style neural-net playground (optimumai playground nn): pick a 2-D dataset, train a tiny MLP, and watch the decision boundary form.

v1.6 adds the Concept Explorer: 30 foundational AI/ML concepts (attention, backprop, gradients, Adam/AdamW, activations, layer norm, PCA, k-means, RL, tokenization, transformer blocks, and more) each rendered as a DAG you step through, with a KaTeX formula and a runnable optimumai code snippet for every step (optimumai explain <concept>), plus a searchable landing page listing all of them (optimumai explore).

Public-API stability

Everything exported from the top-level optimumai namespace (the names in __all__ below) is considered stable and follows semantic versioning: no breaking changes without a major-version bump. Submodule internals and any name prefixed with _ may change between minor releases.

Matrix

A 2-D array with an explainable matrix product.

matmul_trace(other)

Trace C = A @ B, one output cell at a time.

matmul(other, explain=False, level=ExplainLevel.INTERMEDIATE)

Matrix product. Set explain=True to print the cell-by-cell trace.

Vector

A 1-D array with step-by-step explanations of its operations.

dot_trace(other)

Build the full trace of self · other.

dot(other, explain=False, level=ExplainLevel.INTERMEDIATE)

Inner product. Set explain=True to print the full trace.

norm_trace()

Euclidean (L2) norm, ||a|| = √Σ aᵢ².

cosine_similarity_trace(other)

Cosine similarity (a · b) / (||a|| ||b||) — the RAG workhorse.

NTMMemory

A tiny Neural Turing Machine memory bank with content-based addressing.

Wraps :func:ntm_read / :func:ntm_write as stateful operations over a fixed-size memory matrix, mirroring how a real NTM controller would issue a read followed by a write each timestep.

read(key)

Content-addressed read from the current memory state.

write(key, erase, add)

Content-addressed write; updates and returns the new memory state.

Value

A differentiable scalar node in an autograd graph.

zero_grad()

Reset the gradient of every node reachable from this one.

backward()

Run reverse-mode autodiff: seed grad=1 and apply the chain rule.

backward_trace()

Like :meth:backward, but record how each gradient is produced.

For every node processed (output → inputs) we snapshot its children's gradients before and after applying the local _backward, so the delta we log is the exact contribution the chain rule pushed to each child.

backprop(explain=False, level=ExplainLevel.INTERMEDIATE)

Convenience: run backward, optionally rendering the chain-rule trace.

ExplainLevel

Bases: str, Enum

How much detail a rendered trace should surface.

rank property

Ordinal used to gate optional sections (higher shows more).

parse(value) classmethod

Coerce a string (case-insensitive) or enum into an ExplainLevel.

at_least(other)

True when this level is as detailed as other (or more).

Step dataclass

One intermediate move in a computation.

Attributes:

Name Type Description
index int

1-based position in the trace.

title str

Short label, e.g. "Multiply component 0".

expression str

Human-readable computation, e.g. "1 × 4 = 4".

value Any

The numeric result of this step (scalar / vector / matrix).

detail str

Optional longer note shown at higher explain levels.

Trace dataclass

A complete, replayable record of an operation.

Attributes:

Name Type Description
op str

Name of the operation, e.g. "dot" or "attention".

result Any

The final output value.

steps list[Step]

Ordered intermediate steps.

formula str

The closed-form formula, e.g. "a · b = Σ aᵢbᵢ".

why_ai list[str]

Bullet points on where this shows up in real AI systems.

complexity str

Big-O note, surfaced at engineer/researcher levels.

meta dict[str, Any]

Free-form extra data (shapes, hyper-params, ...).

add(title, expression, value=None, detail='')

Append a step and return self for fluent chaining.

render(level=ExplainLevel.INTERMEDIATE, console=None)

Pretty-print this trace to the terminal and return the result.

Imported lazily so the math layer never hard-depends on the UI layer.

last()

Return the final recorded step (raises if the trace is empty).

Course

Navigable view over the ordered curriculum.

tracks()

Lessons grouped by track, preserving order.

next_incomplete(completed)

The first lesson (in course order) not yet in completed.

Lesson dataclass

A single, runnable step in the course.

run(level='intermediate')

Render this lesson's trace at level and return it.

Workbook

Serve and grade compute-the-value exercises (optionally for one lesson).

grade(exercise_id, value)

Grade a submitted value (absolute tolerance, relative fallback for large answers).

KernelWorkbench

Run and grade user-written kernels against a NumPy reference.

reveal(challenge_id)

Return a correct reference solution for the challenge.

submit(challenge_id, user_kernel)

Run user_kernel on the simulator and grade it against the reference.

GpuSim

A serial simulator of a SIMT GPU launch.

Instantiate once per kernel run (or reuse and call :meth:reset). The core is :meth:launch, which iterates every (block, thread) in the requested grid, hands each a :class:ThreadCtx, and tallies memory traffic into self.stats.

reset()

Clear all instrumentation so the simulator can be launched again.

launch(grid, block, kernel, *, flops=0)

Run kernel once per thread in the grid × block launch.

Parameters:

Name Type Description Default
grid int | tuple[int, int]

Blocks in the grid — int (1-D) or (gx, gy) (2-D).

required
block int | tuple[int, int]

Threads per block — int (1-D) or (bx, by) (2-D).

required
kernel Callable[[ThreadCtx], None]

A function taking a single :class:ThreadCtx. It mutates output arrays in place via ctx.gstore (that is where results land).

required
flops int

Total floating-point ops the kernel performs, recorded for arithmetic-intensity reporting. Purely informational.

0

Returns:

Type Description
MemoryStats

(stats, self) — the populated :class:MemoryStats and this sim.

GpuSim

Kernels write their output into caller-owned arrays, so the "output

tuple[MemoryStats, GpuSim]

buffer" is whatever array the caller passed in and the kernel stored to.

Threads are visited block-by-block; within a block they are grouped into warps of :data:WARP_SIZE, and coalescing is judged per warp (consecutive lanes → consecutive addresses). 2-D blocks are flattened in row-major (ty, tx) order for warp grouping, matching CUDA's lane numbering.

KNN

k-nearest-neighbors classifier: store the data, vote at prediction time.

Example

model = KNN(k=1).fit([[0], [1], [10], [11]], [0, 0, 1, 1]) model.predict([[0.5], [10.5]]) array([0, 1])

fit(X, y)

Store the training data (k-NN has no learned parameters).

predict(X)

Classify each row of X by majority vote of its k nearest neighbors.

explain(x_query, level=ExplainLevel.INTERMEDIATE)

Print the full distance/vote trace for one query point and return its label.

PCA

Principal Component Analysis via covariance eigendecomposition.

Example

model = PCA(n_components=1).fit([[0, 0], [1, 1], [2, 2], [3, 3]]) model.transform([[4, 4]]).shape (1, 1)

fit(X)

Compute the mean and top principal components from X.

transform(X)

Project new data onto the fitted principal components.

fit_transform(X)

Fit on X and return the projected training data in one call.

explain(X, level=ExplainLevel.INTERMEDIATE)

Fit, print the full center/covariance/eigen trace, and return the projection.

DecisionTree

A shallow classification tree grown by greedy impurity-reduction splits.

Example

model = DecisionTree(max_depth=1).fit([[0], [1], [10], [11]], [0, 0, 1, 1]) model.predict([[0.5], [10.5]]) array([0, 1])

fit(X, y)

Grow the tree greedily, splitting on maximum information gain.

predict(X)

Route each row of X through the fitted tree to a leaf prediction.

explain(X, y, level=ExplainLevel.INTERMEDIATE)

Fit, print the root-split search trace, and return training predictions.

GaussianNB

Gaussian Naive Bayes classifier.

Example

model = GaussianNB().fit([[0], [1], [10], [11]], [0, 0, 1, 1]) model.predict([[0.5], [10.5]]) array([0, 1])

fit(X, y)

Estimate class priors and per-feature Gaussian parameters.

predict(X)

Classify each row of X by argmax log-posterior.

explain(x_query, level=ExplainLevel.INTERMEDIATE)

Print the full prior/likelihood/posterior trace for one query point.

KMeans

Lloyd's-algorithm k-means clustering.

Example

model = KMeans(k=2).fit([[0], [1], [10], [11]]) model.predict([[0.5], [10.5]]) array([0, 1])

fit(X)

Run Lloyd's algorithm to convergence (or :attr:max_iters).

predict(X)

Assign new points to the nearest fitted centroid.

explain(X, level=ExplainLevel.INTERMEDIATE)

Fit, print the full assign/update trace, and return cluster labels.

LinearRegression

Ordinary least squares via the normal equation.

Example

model = LinearRegression().fit([[1], [2], [3]], [2, 4, 6]) model.predict([[4]]) array([8.])

fit(X, y)

Fit θ by the normal equation and store it on the model.

predict(X)

Predict ŷ = Xθ for new samples (must call :meth:fit first).

explain(X, y, level=ExplainLevel.INTERMEDIATE)

Fit, print the full trace, and return the training predictions.

LogisticRegression

Binary logistic regression trained by batch gradient descent.

Example

model = LogisticRegression(lr=0.5, steps=200).fit([[0], [1], [2], [3]], [0, 0, 1, 1]) model.predict([[0.1], [2.9]]) array([0, 1])

fit(X, y)

Run :attr:steps gradient-descent updates and store the final weights.

predict_proba(X)

Return P(y=1|x) for new samples (must call :meth:fit first).

predict(X)

Return the hard 0/1 label (threshold at 0.5).

explain(X, y, level=ExplainLevel.INTERMEDIATE)

Fit, print the full trace (with :attr:steps gradient updates shown), return labels.

MLP

A stack of dense layers. Hidden layers use activation; output is linear.

Parameters:

Name Type Description Default
n_in int

Number of input features.

required
n_outs Sequence[int]

Sizes of each layer, e.g. [4, 4, 1].

required
activation str

Hidden-layer nonlinearity ("tanh" or "relu").

'tanh'
seed int

Seed for reproducible weight initialization.

0

BPETokenizer

Learn BPE merges from a corpus, then encode new words with them.

Parameters:

Name Type Description Default
num_merges int

How many merge rules to learn during :meth:train.

10

train(corpus)

Learn up to num_merges merge rules from a list of training words.

encode(word)

Apply learned merges (in training order) to tokenize word.

NGramModel

A counting-based n-gram language model with add-k smoothing.

Parameters:

Name Type Description Default
n int

Order of the model (2 = bigram, 3 = trigram, ...). Must be >= 1.

2
k float

Add-k smoothing constant (k=1 is classic Laplace smoothing).

1.0

vocab_size property

Number of distinct tokens seen during :meth:fit (including BOS/EOS).

fit(corpus)

Count n-grams (and their contexts) over a list of sentences.

prob(context, word)

Add-k smoothed P(word | context) for a (n-1)-length context.

sentence_log_prob(sentence)

Sum of log P(w_i | context) over every n-gram in a padded sentence.

perplexity(corpus)

PPL = exp(-(1/N) * sum of log-probs) over every n-gram in corpus.

SkipGramModel

A tiny skip-gram word2vec model trained with full-softmax SGD.

Parameters:

Name Type Description Default
vocab list[str]

The (deduplicated) vocabulary; row i of both embedding matrices belongs to vocab[i].

required
dim int

Embedding dimensionality.

4
seed int

Seed for the initial embedding matrices.

0

predict(center)

Return P(context | center) over the whole vocabulary.

step(center, context, lr=0.1)

One full-softmax SGD step on a single (center, context) pair.

Returns the cross-entropy loss before the update.

most_similar(word, k=3)

Rank the vocabulary by cosine similarity to word in w_in space.

TfidfVectorizer

Fit a vocabulary + idf table on a corpus, then transform documents to vectors.

Mirrors the fit/transform shape of familiar vectorizer APIs, but stays numpy + stdlib only, matching the "tf * smoothed-idf" formula documented in :func:tfidf_trace.

fit(docs)

Learn the vocabulary and idf weights from docs.

transform(docs)

Map each document in docs to a tf-idf row using the fitted idf.

fit_transform(docs)

Fit on docs then transform them in one call.

SGD

Vanilla stochastic gradient descent: w ← w − lr · ∂L/∂w.

Adam

Adam: per-parameter adaptive steps from 1st/2nd gradient moments.

m ← β₁m + (1−β₁)g, v ← β₂v + (1−β₂)g² m̂ = m/(1−β₁ᵗ), v̂ = v/(1−β₂ᵗ) w ← w − lr · m̂ / (√v̂ + ε)

ProgressTracker

Records completed lessons and computes simple progress statistics.

completion_rate(total)

Fraction (0.0–1.0) of total lessons completed.

review_state(lesson_id)

Return the stored SM-2 review state for a lesson, or None.

Quiz

A gradeable quiz over the questions for a single lesson.

questions property

The questions for this lesson, in bank order.

grade(answers)

Grade answers against the correct indices.

A length mismatch is tolerated: any question without a corresponding answer (and any answer past the last question) is scored as wrong.

check(index, choice)

Return whether choice is the correct option for question index.

RAGPipeline

Bases: BaseOp

A minimal, deterministic retrieval-augmented-generation pipeline.

Parameters:

Name Type Description Default
corpus list[str] | None

The documents to retrieve from. Defaults to five ML sentences.

None
dim int

Embedding dimension for the bag-of-hashed-words vectors.

16
seed int

Base seed; each word gets its own vector from seed ⊕ hash(word).

0

trace(query, k=2)

Build the full five-stage RAG trace for query.

forward(query, k=2, explain=False, level='engineer')

Run the pipeline. Set explain=True to print the five-stage trace.

demo() classmethod

A reproducible RAG example over the default ML corpus.

ReviewScheduler

Schedule and surface spaced-repetition reviews over a ProgressTracker.

record(lesson_id, quality, now=None)

Grade a review (0–5), persist the new schedule, and return it.

due(candidates=None, now=None)

Lessons whose review is due (never-reviewed lessons are due immediately).

MDP dataclass

A finite, discounted Markov Decision Process.

Attributes:

Name Type Description
states list[str]

State labels, e.g. ["s0", "s1", "s2"].

actions list[str]

Action labels, e.g. ["left", "right"].

transition ndarray

P[s_idx][a_idx] is a probability vector over next-state indices — transition[s][a][s'] = P(s'|s, a). Must sum to 1 per (s, a).

reward ndarray

R[s_idx][a_idx][s'_idx] = immediate reward for that transition.

gamma float

Discount factor, 0 <= gamma < 1.

terminal frozenset[int]

Optional set of state indices with no outgoing value (their value is fixed at 0 and they are excluded from the max/backup).

expected_backup(values, s, a)

The Bellman backup Σₛ' P(s'|s,a)[R(s,a,s') + γV(s')] for one (s, a).

Attention

Bases: BaseOp

Single-head scaled dot-product attention.

Parameters:

Name Type Description Default
d_k int | None

Key/query dimension used for the 1/√dₖ scaling. If omitted it is inferred from Q at call time.

None

forward(Q, K, V, explain=False, level='intermediate')

Compute attention. Set explain=True to print the four-stage trace.

demo(seed=0) classmethod

A tiny, reproducible 3-token / dₖ=4 example for docs and the CLI.

TransformerBlock

Bases: BaseOp

One pre-norm transformer decoder block (attention + feed-forward).

Parameters:

Name Type Description Default
d_model int

Model / residual-stream dimension.

required
n_heads int

Number of attention heads (must divide d_model).

required
d_ff int | None

Hidden width of the feed-forward network (default 4·d_model).

None

forward(X, explain=False, level='engineer')

Run one transformer block. Set explain=True to print the trace.

demo(seed=0) classmethod

A tiny, reproducible 4-token / d_model=8 / 2-head example.

MultiHeadAttention

Bases: BaseOp

Multi-head scaled dot-product self-attention.

Parameters:

Name Type Description Default
n_heads int

Number of parallel attention heads.

required
d_model int

Model dimension; must be divisible by n_heads.

required

forward(X, causal=False, explain=False, level='engineer')

Compute self-attention. Set explain=True to print the trace.

demo(seed=0) classmethod

A tiny, reproducible 4-token / d_model=8 / 2-head causal example.

TextPipeline

Bases: BaseOp

Push a user's own text through a full, tiny, untrained transformer.

Parameters:

Name Type Description Default
text str

The prompt to run through the pipeline (must be non-empty).

required
layers int

How many :class:TransformerBlock layers to stack.

2
d_model int

Model / residual-stream dimension (must divide by n_heads).

16
n_heads int

Number of attention heads per block.

2
seed int

Seed for the (fixed, untrained) embedding and output-head weights.

0
tokenizer str

"whitespace" (word-level) or "bpe" (via tiktoken).

'whitespace'

forward(explain=False, level='engineer')

Run the pipeline. Set explain=True to print the full stage-by-stage trace.

demo() classmethod

A tiny, reproducible run on a fixed prompt for docs and the CLI.

Tutor

A friendly, optional LLM tutor for the math behind AI.

Safe to construct and call unconditionally: when no client library or key is available, methods return a clear message rather than raising.

available property

True when a client library and an API key are both available.

ask(question, system=None)

Ask a free-form question. Never raises; returns guidance if unavailable.

explain(concept, level=ExplainLevel.INTERMEDIATE)

Explain a math/AI concept at the requested detail level.

explain_trace(trace, level=ExplainLevel.INTERMEDIATE)

Add LLM intuition on top of an OptimumAI :class:Trace.

JEPA

Bases: BaseOp

A minimal Joint-Embedding Predictive Architecture.

Uses a fixed, seeded linear encoder f (shared by context and target) and a linear predictor g living entirely in embedding space. Everything is deterministic given seed so traces are reproducible.

Parameters:

Name Type Description Default
input_dim int

Dimension of a raw input view (e.g. flattened patch).

required
embed_dim int

Dimension of the abstract representation space.

required
seed int

RNG seed for the encoder/predictor weights.

0

forward(context, target, explain=False, level='engineer')

Compute the energy. Set explain=True to print the four-stage trace.

demo(seed=0) classmethod

A tiny, reproducible example: two noisy views of one latent vector.

target = context + small noise mimics two augmentations of the same image, so a well-behaved encoder keeps the energy small.

compare(op_a, op_b, input, explain=False, level=ExplainLevel.ENGINEER)

Compare op_a and op_b on input. Set explain=True to print.

sweep(op, param, values, base_input=None, explain=False, level=ExplainLevel.ENGINEER)

Sweep param of op across values. Set explain=True to print.

adaptive_computation_time(halting_logits, eps=0.01)

Run ACT's halting mechanism over a sequence of ponder-step logits.

Parameters:

Name Type Description Default
halting_logits Iterable[float]

Raw (pre-sigmoid) halting score at each ponder step, in order. Must be non-empty.

required
eps float

Small tolerance; pondering stops once the cumulative halting probability reaches 1 - eps (or the logits run out).

0.01

Returns:

Type Description
dict[str, float | int | ndarray]

A dict with halting_probs (sigmoid of each logit up to and

dict[str, float | int | ndarray]

including the halt step), cumulative (running sum at each step),

dict[str, float | int | ndarray]

halt_step (1-based index of the step that triggered halting),

dict[str, float | int | ndarray]

remainder (weight given to the final step), and ponder_cost

dict[str, float | int | ndarray]

(halt_step + remainder).

attention_read(query, memory)

Read a weighted blend of memory rows using content-based attention.

Parameters:

Name Type Description Default
query ndarray

A 1-D vector of shape (d,) — "what am I looking for?".

required
memory ndarray

A 2-D array of shape (n, d) — n memory slots of dimension d — "what is stored?".

required

Returns:

Type Description
ndarray

The blended read vector, shape (d,).

ntm_read(key, memory, beta=1.0)

Content-addressed read: blend memory rows by cosine similarity to key.

Parameters:

Name Type Description Default
key ndarray

1-D vector of shape (d,) used to address memory.

required
memory ndarray

2-D array of shape (n, d) — the memory bank.

required
beta float

Sharpness ("key strength"); larger values make addressing more peaked around the best-matching slot.

1.0

Returns:

Type Description
ndarray

The read vector, shape (d,).

ntm_write(memory, key, erase, add, beta=1.0)

Content-addressed write: erase then add, blended by addressing weights.

Mᵢ ← Mᵢ · (1 − wᵢ · erase) + wᵢ · add for every slot i.

Parameters:

Name Type Description Default
memory ndarray

2-D array of shape (n, d) — the memory bank before the write.

required
key ndarray

1-D vector of shape (d,) used to address memory.

required
erase ndarray

1-D vector of shape (d,) in [0, 1] — what to clear.

required
add ndarray

1-D vector of shape (d,) — what to write in.

required
beta float

Addressing sharpness, as in :func:ntm_read.

1.0

Returns:

Type Description
ndarray

The new memory bank, shape (n, d).

derivative(f, x, h=1e-05, explain=False, level=ExplainLevel.INTERMEDIATE)

Numeric derivative of f at x (central difference).

gradient(f, point, h=1e-05, explain=False, level=ExplainLevel.INTERMEDIATE)

Numeric gradient of f at point.

forward_diffusion(x0, timesteps=10, beta_start=0.0001, beta_end=0.2, seed=0, explain=False, level=ExplainLevel.INTERMEDIATE)

Noise x0 to timestep T. Set explain=True to print the trace.

embedding_lookup(tokens, dim=4, seed=0, explain=False, level=ExplainLevel.INTERMEDIATE)

Look up embedding rows for tokens. Set explain=True to print it.

nearest_neighbors(query, vocab, dim=4, seed=0, k=3, explain=False, level=ExplainLevel.INTERMEDIATE)

Return the top-k cosine neighbours of query within vocab.

bleu(candidate, reference, max_n=4)

Return the BLEU score of candidate against reference in [0, 1].

ece(confidences, correct, n_bins=5)

Return the Expected Calibration Error in [0, 1] (lower is better calibrated).

exact_match(candidate, reference)

1.0 if the normalized token sequences are identical, else 0.0 (SQuAD-style).

faithfulness_score(answer, context)

Return the claim-overlap faithfulness heuristic in [0, 1] (see module docstring).

rouge_l(candidate, reference)

Return the ROUGE-L F-measure in [0, 1].

rouge_n(candidate, reference, n=1)

Return ROUGE-N recall (the conventionally reported value) in [0, 1].

token_f1(candidate, reference)

Return the token-level F1 of candidate against reference in [0, 1].

kv_cache_size(n_layers, n_heads, head_dim, seq_len, batch=1, bytes_per_elem=2, kv_heads=None)

Return the KV cache size in bytes (no trace).

kv_heads defaults to n_heads (Multi-Head Attention); pass a smaller number for Grouped-Query Attention or 1 for Multi-Query Attention.

integrate(f, a, b, method='trapezoid', n=100, explain=False, level=ExplainLevel.INTERMEDIATE)

Return ∫ f from a to b. Set explain=True to print the trace.

vram_estimate(params_billions, precision_bytes=2, training=True, optimizer='adam', batch=1, seq_len=2048, hidden=4096, n_layers=32, activation_factor=None)

Return the estimated total VRAM in GB (no trace).

flash_attention(Q, K, V, block_size=2, explain=False, level=ExplainLevel.ENGINEER)

Return flash attention output. Set explain=True to print the trace.

lora(d_in=8, d_out=8, rank=2, seed=0, explain=False, level=ExplainLevel.ENGINEER)

Return the LoRA parameter reduction factor. explain=True prints the trace.

quantize(x, bits=8, scheme='symmetric', granularity='tensor')

Quantize x to integer codes.

Returns:

Type Description

(q, scale, zero_point) where q are the clipped integer codes and

scale/zero_point are the params needed to dequantize.

dpo(prompt='Explain gravity.', beta=0.1, seed=0, explain=False, level=ExplainLevel.ENGINEER)

Return the DPO loss for one preference triple. explain=True prints the trace.

run_repl()

Start the interactive read-eval-print loop.

superposition(n_features, n_neurons, sparsity=0.7, seed=0, explain=False, level=ExplainLevel.ENGINEER)

Run the toy superposition model, returning the recovered vector ĥ.

Set explain=True to print the encode/interference/decode trace.

generate(prompt, provider='auto', model=None, max_tokens=64, temperature=0.7)

Generate a continuation of prompt and return the text.

provider="auto" picks the best available (Ollama → HF → Anthropic → toy).

minimize(loss_fn, params, optimizer, steps=50, explain=False, level=ExplainLevel.INTERMEDIATE)

Minimize loss_fn over params and return the final parameter values.

softmax(x, temperature=1.0, explain=False, level=ExplainLevel.INTERMEDIATE)

Return the softmax of x. Set explain=True to print the trace.

softmax_trace(x, temperature=1.0)

Build the full trace of softmax(x) with an optional temperature.

policy_iteration(mdp, theta=1e-06, max_iterations=1000, explain=False, level=ExplainLevel.ENGINEER)

Return V^π* for mdp. explain=True prints the full trace.

ppo_clip(batch=None, epsilon=0.2, explain=False, level=ExplainLevel.ENGINEER)

Return the PPO clipped loss for batch. explain=True prints the trace.

reinforce(env=None, episodes=200, alpha=0.1, gamma=0.99, seed=0, explain=False, level=ExplainLevel.ENGINEER)

Return the learned action-probability vector. explain=True prints the trace.

sarsa(world=None, episodes=500, alpha=0.5, gamma=0.9, epsilon=0.1, seed=0, explain=False, level=ExplainLevel.ENGINEER)

Return the learned Q-table. explain=True prints the full trace.

value_iteration(mdp, theta=1e-06, max_iterations=1000, explain=False, level=ExplainLevel.ENGINEER)

Return V* for mdp. explain=True prints the full trace.

alpha_beta(node, maximizing=True, alpha=float('-inf'), beta=float('inf'))

Return the minimax value of node's game tree, computed with pruning.

astar(problem, start, goal)

Return the optimal path from start to goal via A* (admissible h assumed).

bfs(problem, start, goal)

Return the fewest-edges path from start to goal (BFS).

dfs(problem, start, goal)

Return a path from start to goal found via DFS (not necessarily optimal).

greedy_best_first(problem, start, goal)

Return a path from start to goal via greedy best-first search.

minimax(node, maximizing=True)

Return the minimax value of the root of node's game tree.

Return the minimum-cost path from start to goal (UCS / Dijkstra).

differentiate(expression, var='x', at=None, explain=False, level=ExplainLevel.INTERMEDIATE)

Differentiate expression w.r.t. var. Set explain=True to print.

positional_encoding(seq_len, d_model, explain=False, level=ExplainLevel.INTERMEDIATE)

Return the sinusoidal PE matrix. Set explain=True to print the trace.

get_tutorial(name)

Load and build the tutorial called name (lazy import).

list_tutorials()

The available tutorial names.

avg_pool2d(x, kernel_size=2, stride=None)

Downsample x by averaging each kernel_size×kernel_size window.

cnn_forward(image, kernel, dense_weights, dense_bias, pool_size=2)

Run one image through conv -> ReLU -> max-pool -> flatten -> dense -> softmax.

conv2d(x, w, stride=1, padding=0, mode='cross-correlate')

Slide kernel w over image x and return the output feature map.

mode="cross-correlate" (default) is what every deep learning framework calls "convolution": no kernel flip. mode="convolve" flips the kernel 180° first, matching the classical signal-processing definition.

max_pool2d(x, kernel_size=2, stride=None)

Downsample x by taking the max of each kernel_size×kernel_size window.

sobel_edges(x, padding=1)

Return (magnitude, orientation) gradient maps for image x.

orientation is in radians, from atan2(Gy, Gx) (range (-π, π]).

optix_js() cached

Return the OptiX bundle source (window.OptiX IIFE).

Raises FileNotFoundError if the compiled asset wasn't shipped — build it with cd web && npm install && npm run build.

render_concept(name, fmt='png', out=None)

Render name as fmt ("png" or "gif") to out (or a default path).

explain(concept, out=None, open_browser=True)

Build an interactive DAG explainer (formula + code per step) for concept.

Parameters:

Name Type Description Default
concept str

One of :func:list_explain_concepts (case/dash/space-insensitive).

required
out str | None

Output HTML path; defaults to explain_{concept}.html.

None
open_browser bool

Open the generated file in the default browser.

True

Returns:

Type Description
str

The path the HTML was written to.

explore_concepts(out=None, open_browser=True)

Build a searchable landing page linking to every :func:explain concept.

Each card links to explain_{key}.html alongside the generated page; run :func:explain for a concept (or optimumai explain <concept>) to inspect one directly.

list_explain_concepts()

Return every concept key :func:explain accepts, sorted alphabetically.

editable_plot(expression='a*x^2 + b*x + c', params=None, xrange=(-5.0, 5.0), out=None)

Build the interactive editor HTML.

If out is given, write the file and return its path; otherwise return the HTML string. Parameters (single-letter names other than x) get sliders unless you pass an explicit params mapping.

nn_playground(out=None)

A TensorFlow-Playground-style neural-net playground, powered by OptiX.

Pick a 2-D dataset (XOR / circle / spiral), set the learning rate and hidden width, and train a tiny MLP while its decision boundary forms live. The math — seeded MLP init, forward pass, and backprop — is OptiX, a typed, unit-tested TypeScript kit compiled into OptimumAI. Self-contained, offline.

playground(name, out=None)

Build the named playground ("attention", "kmeans", "astar", or "nn").

describe(data)

Numpy summary statistics for a batch of numbers.

Returns n, mean, std, min, q1, median, q3, max. Raises :class:ValueError on empty data.

plot_code(data, kind='bar', **opts)

Return the exact, copy-paste-runnable matplotlib+numpy source.

The string always starts with import numpy as np and import matplotlib.pyplot as plt, builds the same array Plot Studio used, calls the matching plotting function, labels the axes/title, and ends with plt.show().

kind is one of "bar", "hist", "scatter", "box", "line", "pie", "violin". For "scatter", data may be a list of (x, y) pairs or two equal-length sequences [xs, ys]. opts accepts title (all kinds) and bins ("hist" only, default 10).

Raises :class:ValueError on empty data or an unknown kind.

plot_data(data, kind='bar', *, out='plot.png', **opts)

Render data as kind with matplotlib and save it to out.

Returns the path to the saved PNG. See :func:plot_code for the supported kind values, the accepted data shapes, and opts.

plot_studio_playground(out=None)

Build the Plot Studio HTML playground: type numbers, get chart + code.

A textarea accepts numbers separated by commas, whitespace, or newlines (for scatter, one x,y pair per line). A dropdown picks one of the seven supported chart kinds. Three panels update live as you type: the rendered chart (drawn on a <canvas> by hand-rolled JavaScript), a numpy-style summary-statistics table, and the exact matplotlib+numpy source that reproduces the current chart, with a Copy button.

Fully self-contained: inline CSS and vanilla JavaScript, no CDN, no server, works offline. Seeded with example data so it renders on open.


Submodule reference

optimumai.algebra

Linear algebra with step-by-step explanations.

Matrix

A 2-D array with an explainable matrix product.

matmul_trace(other)

Trace C = A @ B, one output cell at a time.

matmul(other, explain=False, level=ExplainLevel.INTERMEDIATE)

Matrix product. Set explain=True to print the cell-by-cell trace.

Vector

A 1-D array with step-by-step explanations of its operations.

dot_trace(other)

Build the full trace of self · other.

dot(other, explain=False, level=ExplainLevel.INTERMEDIATE)

Inner product. Set explain=True to print the full trace.

norm_trace()

Euclidean (L2) norm, ||a|| = √Σ aᵢ².

cosine_similarity_trace(other)

Cosine similarity (a · b) / (||a|| ||b||) — the RAG workhorse.

optimumai.probability

Probability building blocks used across modern AI.

softmax_trace(x, temperature=1.0)

Build the full trace of softmax(x) with an optional temperature.

optimumai.calculus

Calculus: numeric derivatives, gradients, and the chain rule.

chain_rule_trace(x=1.5)

Demonstrate the chain rule on f(x) = tanh(x²) — exact vs numeric.

dy/dx = tanh'(x²) · d(x²)/dx = (1 − tanh²(x²)) · 2x.

derivative_trace(f, x, h=1e-05, label='f')

Central-difference derivative of a scalar function at x.

gradient(f, point, h=1e-05, explain=False, level=ExplainLevel.INTERMEDIATE)

Numeric gradient of f at point.

gradient_trace(f, point, h=1e-05)

Numeric gradient (vector of partial derivatives) of f at point.

optimumai.autograd

A tiny scalar autograd engine (micrograd-style) with a traceable backward pass.

Value

A differentiable scalar node in an autograd graph.

zero_grad()

Reset the gradient of every node reachable from this one.

backward()

Run reverse-mode autodiff: seed grad=1 and apply the chain rule.

backward_trace()

Like :meth:backward, but record how each gradient is produced.

For every node processed (output → inputs) we snapshot its children's gradients before and after applying the local _backward, so the delta we log is the exact contribution the chain rule pushed to each child.

backprop(explain=False, level=ExplainLevel.INTERMEDIATE)

Convenience: run backward, optionally rendering the chain-rule trace.

optimumai.optimization

Optimizers and the training loop that turns gradients into learning.

SGD

Vanilla stochastic gradient descent: w ← w − lr · ∂L/∂w.

Adam

Adam: per-parameter adaptive steps from 1st/2nd gradient moments.

m ← β₁m + (1−β₁)g, v ← β₂v + (1−β₂)g² m̂ = m/(1−β₁ᵗ), v̂ = v/(1−β₂ᵗ) w ← w − lr · m̂ / (√v̂ + ε)

descent_demo(optimizer='adam', steps=50)

Minimize the bowl f(x, y) = (x−3)² + (y+1)² from the origin.

The global minimum is at (3, −1) with loss 0 — an easy target to watch the optimizer walk toward.

minimize(loss_fn, params, optimizer, steps=50, explain=False, level=ExplainLevel.INTERMEDIATE)

Minimize loss_fn over params and return the final parameter values.

minimize_trace(loss_fn, params, optimizer, steps=50)

Run steps optimization iterations, recording the loss each time.

loss_fn must rebuild a scalar-:class:Value loss from the (persistent) params on every call. Each iteration recomputes the graph, backprops, and lets the optimizer nudge the parameters.

optimumai.neural_networks

Neural networks built on the scalar autograd engine.

MLP

A stack of dense layers. Hidden layers use activation; output is linear.

Parameters:

Name Type Description Default
n_in int

Number of input features.

required
n_outs Sequence[int]

Sizes of each layer, e.g. [4, 4, 1].

required
activation str

Hidden-layer nonlinearity ("tanh" or "relu").

'tanh'
seed int

Seed for reproducible weight initialization.

0

Layer

A dense layer: n_out neurons, each seeing all inputs.

Neuron

A single neuron: activation(Σ wᵢxᵢ + b).

forward_backward_trace(mlp, x, target)

One example: forward to a prediction, squared-error loss, then backprop.

train(steps=100, lr=0.05, seed=0, explain=False, level=ExplainLevel.INTERMEDIATE)

Run the training demo; returns the final parameter vector.

train_demo(steps=100, lr=0.05, seed=0)

Train a 3→4→4→1 MLP on the toy set and watch the loss fall.

optimumai.transformers

Transformer math, explained one stage at a time.

Attention

Bases: BaseOp

Single-head scaled dot-product attention.

Parameters:

Name Type Description Default
d_k int | None

Key/query dimension used for the 1/√dₖ scaling. If omitted it is inferred from Q at call time.

None

forward(Q, K, V, explain=False, level='intermediate')

Compute attention. Set explain=True to print the four-stage trace.

demo(seed=0) classmethod

A tiny, reproducible 3-token / dₖ=4 example for docs and the CLI.

TransformerBlock

Bases: BaseOp

One pre-norm transformer decoder block (attention + feed-forward).

Parameters:

Name Type Description Default
d_model int

Model / residual-stream dimension.

required
n_heads int

Number of attention heads (must divide d_model).

required
d_ff int | None

Hidden width of the feed-forward network (default 4·d_model).

None

forward(X, explain=False, level='engineer')

Run one transformer block. Set explain=True to print the trace.

demo(seed=0) classmethod

A tiny, reproducible 4-token / d_model=8 / 2-head example.

MultiHeadAttention

Bases: BaseOp

Multi-head scaled dot-product self-attention.

Parameters:

Name Type Description Default
n_heads int

Number of parallel attention heads.

required
d_model int

Model dimension; must be divisible by n_heads.

required

forward(X, causal=False, explain=False, level='engineer')

Compute self-attention. Set explain=True to print the trace.

demo(seed=0) classmethod

A tiny, reproducible 4-token / d_model=8 / 2-head causal example.

positional_encoding(seq_len, d_model, explain=False, level=ExplainLevel.INTERMEDIATE)

Return the sinusoidal PE matrix. Set explain=True to print the trace.

positional_encoding_trace(seq_len, d_model)

Build the full trace of the sinusoidal positional encoding matrix.

optimumai.world_models

World models — learning to predict in representation space (LeCun's JEPA).

JEPA

Bases: BaseOp

A minimal Joint-Embedding Predictive Architecture.

Uses a fixed, seeded linear encoder f (shared by context and target) and a linear predictor g living entirely in embedding space. Everything is deterministic given seed so traces are reproducible.

Parameters:

Name Type Description Default
input_dim int

Dimension of a raw input view (e.g. flattened patch).

required
embed_dim int

Dimension of the abstract representation space.

required
seed int

RNG seed for the encoder/predictor weights.

0

forward(context, target, explain=False, level='engineer')

Compute the energy. Set explain=True to print the four-stage trace.

demo(seed=0) classmethod

A tiny, reproducible example: two noisy views of one latent vector.

target = context + small noise mimics two augmentations of the same image, so a well-behaved encoder keeps the energy small.

optimumai.interpretability

Interpretability — why neurons are polysemantic (Anthropic's superposition).

superposition_trace(n_features, n_neurons, sparsity=0.7, seed=0)

Build the full trace of a toy superposition encode/decode cycle.

Parameters:

Name Type Description Default
n_features int

Number of distinct features (must exceed n_neurons).

required
n_neurons int

Number of neurons / activation dimensions.

required
sparsity float

Fraction of features that are inactive on a given input.

0.7
seed int

RNG seed for the weight matrix and the sparse feature vector.

0

optimumai.frontier

Frontier concepts — how today's large models are actually built and run.

FlashAttention (IO-aware tiling + online softmax), quantization (int8/int4), LoRA (parameter-efficient fine-tuning), and DPO (preference alignment).

flash_attention_trace(Q, K, V, block_size=2)

Build the full trace of tiled attention with online softmax.

Parameters:

Name Type Description Default
Q

Query matrix (n_q × d).

required
K

Key matrix (n_k × d).

required
V

Value matrix (n_k × d_v).

required
block_size int

How many rows of K/V (and of Q) live in one SRAM tile.

2

lora_trace(d_in=8, d_out=8, rank=2, seed=0)

Build the full trace of a LoRA adapter over a frozen weight W₀.

Constructs a seeded frozen W₀ plus the LoRA factors A (Gaussian) and B (zeros), shows that ΔW = B·A = 0 at init, counts the trainable parameters saved, and demonstrates the forward pass y = (W₀ + B·A)·x.

dequantize(q, scale, zero_point)

Rebuild approximate floats: x̂ = (q − zero_point)·scale.

quantize(x, bits=8, scheme='symmetric', granularity='tensor')

Quantize x to integer codes.

Returns:

Type Description

(q, scale, zero_point) where q are the clipped integer codes and

scale/zero_point are the params needed to dequantize.

quantize_trace(x, bits=8, scheme='symmetric', granularity='tensor')

Build the full trace of quantizing (and dequantizing) x.

dpo(prompt='Explain gravity.', beta=0.1, seed=0, explain=False, level=ExplainLevel.ENGINEER)

Return the DPO loss for one preference triple. explain=True prints the trace.

dpo_trace(prompt='Explain gravity.', beta=0.1, seed=0)

Build the full trace of the DPO loss on one toy preference triple.

Uses small seeded per-token log-probabilities for a policy model and a frozen reference model over a chosen and a rejected response, then walks through the implicit rewards, the preference margin, and the final DPO loss.

optimumai.embeddings

Embeddings — turning discrete tokens into dense vectors.

embedding_lookup(tokens, dim=4, seed=0, explain=False, level=ExplainLevel.INTERMEDIATE)

Look up embedding rows for tokens. Set explain=True to print it.

embedding_lookup_trace(tokens, dim=4, seed=0)

Build the full trace of looking up embedding rows for tokens.

nearest_neighbors(query, vocab, dim=4, seed=0, k=3, explain=False, level=ExplainLevel.INTERMEDIATE)

Return the top-k cosine neighbours of query within vocab.

nearest_neighbors_trace(query, vocab, dim=4, seed=0, k=3)

Rank vocab by cosine similarity to query in embedding space.

optimumai.rag

Retrieval-augmented generation, traced end to end.

RAGPipeline

Bases: BaseOp

A minimal, deterministic retrieval-augmented-generation pipeline.

Parameters:

Name Type Description Default
corpus list[str] | None

The documents to retrieve from. Defaults to five ML sentences.

None
dim int

Embedding dimension for the bag-of-hashed-words vectors.

16
seed int

Base seed; each word gets its own vector from seed ⊕ hash(word).

0

trace(query, k=2)

Build the full five-stage RAG trace for query.

forward(query, k=2, explain=False, level='engineer')

Run the pipeline. Set explain=True to print the five-stage trace.

demo() classmethod

A reproducible RAG example over the default ML corpus.

rag_flow(query=_DEFAULT_QUERY, k=2, out='rag_explainer.html')

Build a RAG :class:~optimumai.core.flow_trace.FlowTrace and render it to a self-contained D3 + KaTeX HTML file.

Parameters

query: The question to retrieve context for. Cosine scores are computed from the real :class:~optimumai.rag.pipeline.RAGPipeline — not toy values. k: Top-k chunks to retrieve. out: Output HTML path. Defaults to "rag_explainer.html" in the current directory.

Returns

str The path to the generated HTML file.

Examples

CLI::

optimumai flow rag
optimumai flow rag --out my_rag.html

Python::

from optimumai.rag.flow import rag_flow
rag_flow(query="How tall is the Eiffel Tower?", out="rag.html")

build_rag_trace(query=_DEFAULT_QUERY, k=2, rerank_noise=0.04, seed=0)

Build a :class:FlowTrace for a RAG pipeline run.

Parameters

query: The user question to retrieve context for. k: Top-k chunks to retrieve. rerank_noise: Small perturbation added to cosine scores to simulate a cross-encoder reranker (real re-ranking shifts scores without changing the embedding). seed: Random seed for the embedding model (for reproducibility).

Returns

FlowTrace A validated trace with real cosine similarity scores computed by the actual :class:~optimumai.rag.pipeline.RAGPipeline.

optimumai.diffusion

Diffusion — forward noising and the reverse denoising idea.

forward_diffusion(x0, timesteps=10, beta_start=0.0001, beta_end=0.2, seed=0, explain=False, level=ExplainLevel.INTERMEDIATE)

Noise x0 to timestep T. Set explain=True to print the trace.

forward_diffusion_trace(x0, timesteps=10, beta_start=0.0001, beta_end=0.2, seed=0)

Build the full trace of DDPM forward diffusion applied to x0.

optimumai.foundations

Foundations of the AI stack — the math, frameworks, and hardware underneath.

Tensors & integration, the PyTorch and JAX programming models (autograd, grad/jit/vmap), and the systems layer: the CUDA execution & memory model, tiled matmul kernels, the KV cache, and a VRAM budget calculator.

tiled_matmul(M=64, N=64, K=64, tile=16, explain=False, level=ExplainLevel.ENGINEER)

Return C = A·B (seeded). Set explain=True to print the tiling trace.

tiled_matmul_trace(M=64, N=64, K=64, tile=16)

Trace the naive → tiled → coalesced optimization of C = A · B.

Parameters:

Name Type Description Default
M, N, K

Dimensions of C = A·B where A is M×K and B is K×N.

required
tile int

Side length of the square tile loaded into shared memory. The global-memory reads drop by roughly this factor.

16

The result is the real product matrix C (computed with a seeded NumPy A @ B). meta carries naive_reads, tiled_reads and reduction_factor (which equals tile).

memory_hierarchy_trace()

Trace the GPU memory hierarchy from fastest/smallest to slowest/largest.

Walks registers → shared memory (L1) → global memory (VRAM) with approximate latencies and bandwidths, then explains arithmetic intensity (FLOP per byte) and why most kernels are memory-bound rather than compute-bound. The result is a small dict summarizing each level.

thread_hierarchy_trace(grid_dim=3, block_dim=4, thread_id=None)

Trace the grid → block → warp → thread hierarchy of a GPU launch.

Parameters:

Name Type Description Default
grid_dim int

Number of blocks in the (1-D) grid.

3
block_dim int

Number of threads per block.

4
thread_id int | None

A global thread index to decompose back into (block_idx, thread_idx). Defaults to the middle thread. If given, the trace's result is that (block_idx, thread_idx) tuple; otherwise the result is the total thread count.

None

grad_trace(f, x, h=1e-05)

Show JAX's grad idea via a central-difference derivative of f at x.

We approximate f'(x) numerically here so the concept is visible without a tracer. JAX computes the exact derivative by tracing the pure function f and running reverse-mode autodiff over that trace — same answer, no round-off.

pytree_trace(tree)

Flatten a nested dict/list/tuple tree to leaves + structure, then rebuild.

Every JAX transformation operates over pytrees: model parameters are nested dicts, optimizer state is a pytree, grad returns a pytree matching the inputs. Flatten → transform the flat leaves → unflatten is the pattern.

vmap_trace(f, batch)

Demonstrate vmap: apply pure f across a batch axis with no Python loop.

We compute the loop version and the stacked version and show they agree. JAX's vmap does the same thing by pushing a batch dimension through the traced function, letting XLA vectorize it — no interpreter loop at runtime.

kv_cache_size(n_layers, n_heads, head_dim, seq_len, batch=1, bytes_per_elem=2, kv_heads=None)

Return the KV cache size in bytes (no trace).

kv_heads defaults to n_heads (Multi-Head Attention); pass a smaller number for Grouped-Query Attention or 1 for Multi-Query Attention.

kv_cache_trace(n_layers, n_heads, head_dim, seq_len, batch=1, bytes_per_elem=2, kv_heads=None)

Build the full trace of a transformer KV cache's memory footprint.

integrate(f, a, b, method='trapezoid', n=100, explain=False, level=ExplainLevel.INTERMEDIATE)

Return ∫ f from a to b. Set explain=True to print the trace.

integrate_trace(f, a, b, method='trapezoid', n=100)

Estimate ∫ f(x) dx from a to b and trace the steps.

Parameters:

Name Type Description Default
f Callable[[ndarray], ndarray]

A vectorized function accepting a NumPy array and returning one.

required
a float

Lower bound of integration.

required
b float

Upper bound of integration.

required
method str

"trapezoid" (grid of trapezoids) or "monte_carlo" (average of f at seeded random points, times the width).

'trapezoid'
n int

Number of grid intervals (trapezoid) or samples (Monte Carlo).

100

tensor_intro_trace()

Build a trace introducing tensors as n-dimensional arrays.

pytorch_autograd(explain=False, level=ExplainLevel.ENGINEER)

Return the loss of the demo neuron. Set explain=True to print the trace.

pytorch_autograd_trace()

Build a tiny neuron with :class:Value and narrate the PyTorch equivalents.

We compute y = tanh(w*x + b) and loss = (y - target)**2, then call loss.backward(). Every step maps back to a line of PyTorch so you can see that torch.autograd is just reverse-mode autodiff over a dynamic graph.

vram_estimate(params_billions, precision_bytes=2, training=True, optimizer='adam', batch=1, seq_len=2048, hidden=4096, n_layers=32, activation_factor=None)

Return the estimated total VRAM in GB (no trace).

vram_trace(params_billions, precision_bytes=2, training=True, optimizer='adam', batch=1, seq_len=2048, hidden=4096, n_layers=32, activation_factor=None)

Build the full trace of an LLM's VRAM budget, component by component.

optimumai.kernels

GPU kernels from scratch — a pure-Python simulator, real kernels, and challenges.

KernelWorkbench

Run and grade user-written kernels against a NumPy reference.

reveal(challenge_id)

Return a correct reference solution for the challenge.

submit(challenge_id, user_kernel)

Run user_kernel on the simulator and grade it against the reference.

GpuSim

A serial simulator of a SIMT GPU launch.

Instantiate once per kernel run (or reuse and call :meth:reset). The core is :meth:launch, which iterates every (block, thread) in the requested grid, hands each a :class:ThreadCtx, and tallies memory traffic into self.stats.

reset()

Clear all instrumentation so the simulator can be launched again.

launch(grid, block, kernel, *, flops=0)

Run kernel once per thread in the grid × block launch.

Parameters:

Name Type Description Default
grid int | tuple[int, int]

Blocks in the grid — int (1-D) or (gx, gy) (2-D).

required
block int | tuple[int, int]

Threads per block — int (1-D) or (bx, by) (2-D).

required
kernel Callable[[ThreadCtx], None]

A function taking a single :class:ThreadCtx. It mutates output arrays in place via ctx.gstore (that is where results land).

required
flops int

Total floating-point ops the kernel performs, recorded for arithmetic-intensity reporting. Purely informational.

0

Returns:

Type Description
MemoryStats

(stats, self) — the populated :class:MemoryStats and this sim.

GpuSim

Kernels write their output into caller-owned arrays, so the "output

tuple[MemoryStats, GpuSim]

buffer" is whatever array the caller passed in and the kernel stored to.

Threads are visited block-by-block; within a block they are grouped into warps of :data:WARP_SIZE, and coalescing is judged per warp (consecutive lanes → consecutive addresses). 2-D blocks are flattened in row-major (ty, tx) order for warp grouping, matching CUDA's lane numbering.

MemoryStats dataclass

Instrumentation gathered over one :meth:GpuSim.launch.

Attributes:

Name Type Description
global_reads int

Number of scalar reads from global memory (VRAM).

global_writes int

Number of scalar writes to global memory.

shared_reads int

Number of scalar reads from shared memory (SRAM).

shared_writes int

Number of scalar writes to shared memory.

flops int

Floating-point operations declared for the launch (for intensity).

coalesced bool | None

Whether every warp's global accesses hit consecutive addresses. None means no global accesses were recorded.

global_bytes property

Total bytes moved to/from global memory (fp32 accounting).

total_global property

Total scalar global-memory accesses (reads + writes).

arithmetic_intensity()

FLOPs per global byte moved — the roofline x-axis.

Higher means more compute-bound (data is reused); lower means memory-bound (the ALUs starve waiting on VRAM). Returns 0.0 when no bytes moved.

as_dict()

A JSON-friendly summary suitable for a :class:Trace meta block.

ThreadCtx

The per-thread handle a kernel uses to read/write memory and see its index.

A fresh context is created for every simulated thread by :meth:GpuSim.launch; kernels never construct one directly. All memory helpers funnel through the parent :class:GpuSim so a single :class:MemoryStats accumulates across the whole grid.

gload(array, index)

Read array[index] from global memory, counting the access.

gstore(array, index, value)

Write value to array[index] in global memory, counting the access.

sload(buffer, index)

Read buffer[index] from shared memory, counting the access.

sstore(buffer, index, value)

Write value to buffer[index] in shared memory, counting the access.

available_backends() cached

Backends usable on this machine. Always includes "simulator".

backend_report()

A human-readable summary of what's available and how to enable more.

list_kernels()

Names of the available kernels, simple → complex.

run_kernel(name, explain=False, level='engineer')

Run a kernel by name; explain=True renders the trace.

optimumai.ml

Classical machine learning — the algorithms that predate (and still outlive) deep nets.

Every estimator here follows the same shape as the rest of OptimumAI: a <name>_trace(...) function builds a full step-by-step :class:~optimumai.core.trace.Trace with real numbers, and a thin class (LinearRegression, KMeans, ...) wraps it in the familiar fit/predict interface.

DecisionTree

A shallow classification tree grown by greedy impurity-reduction splits.

Example

model = DecisionTree(max_depth=1).fit([[0], [1], [10], [11]], [0, 0, 1, 1]) model.predict([[0.5], [10.5]]) array([0, 1])

fit(X, y)

Grow the tree greedily, splitting on maximum information gain.

predict(X)

Route each row of X through the fitted tree to a leaf prediction.

explain(X, y, level=ExplainLevel.INTERMEDIATE)

Fit, print the root-split search trace, and return training predictions.

KMeans

Lloyd's-algorithm k-means clustering.

Example

model = KMeans(k=2).fit([[0], [1], [10], [11]]) model.predict([[0.5], [10.5]]) array([0, 1])

fit(X)

Run Lloyd's algorithm to convergence (or :attr:max_iters).

predict(X)

Assign new points to the nearest fitted centroid.

explain(X, level=ExplainLevel.INTERMEDIATE)

Fit, print the full assign/update trace, and return cluster labels.

KNN

k-nearest-neighbors classifier: store the data, vote at prediction time.

Example

model = KNN(k=1).fit([[0], [1], [10], [11]], [0, 0, 1, 1]) model.predict([[0.5], [10.5]]) array([0, 1])

fit(X, y)

Store the training data (k-NN has no learned parameters).

predict(X)

Classify each row of X by majority vote of its k nearest neighbors.

explain(x_query, level=ExplainLevel.INTERMEDIATE)

Print the full distance/vote trace for one query point and return its label.

LinearRegression

Ordinary least squares via the normal equation.

Example

model = LinearRegression().fit([[1], [2], [3]], [2, 4, 6]) model.predict([[4]]) array([8.])

fit(X, y)

Fit θ by the normal equation and store it on the model.

predict(X)

Predict ŷ = Xθ for new samples (must call :meth:fit first).

explain(X, y, level=ExplainLevel.INTERMEDIATE)

Fit, print the full trace, and return the training predictions.

LogisticRegression

Binary logistic regression trained by batch gradient descent.

Example

model = LogisticRegression(lr=0.5, steps=200).fit([[0], [1], [2], [3]], [0, 0, 1, 1]) model.predict([[0.1], [2.9]]) array([0, 1])

fit(X, y)

Run :attr:steps gradient-descent updates and store the final weights.

predict_proba(X)

Return P(y=1|x) for new samples (must call :meth:fit first).

predict(X)

Return the hard 0/1 label (threshold at 0.5).

explain(X, y, level=ExplainLevel.INTERMEDIATE)

Fit, print the full trace (with :attr:steps gradient updates shown), return labels.

GaussianNB

Gaussian Naive Bayes classifier.

Example

model = GaussianNB().fit([[0], [1], [10], [11]], [0, 0, 1, 1]) model.predict([[0.5], [10.5]]) array([0, 1])

fit(X, y)

Estimate class priors and per-feature Gaussian parameters.

predict(X)

Classify each row of X by argmax log-posterior.

explain(x_query, level=ExplainLevel.INTERMEDIATE)

Print the full prior/likelihood/posterior trace for one query point.

PCA

Principal Component Analysis via covariance eigendecomposition.

Example

model = PCA(n_components=1).fit([[0, 0], [1, 1], [2, 2], [3, 3]]) model.transform([[4, 4]]).shape (1, 1)

fit(X)

Compute the mean and top principal components from X.

transform(X)

Project new data onto the fitted principal components.

fit_transform(X)

Fit on X and return the projected training data in one call.

explain(X, level=ExplainLevel.INTERMEDIATE)

Fit, print the full center/covariance/eigen trace, and return the projection.

decision_tree_trace(X, y, max_depth=2, criterion='gini')

Build a trace of the best-split search at the root, then grow a shallow tree.

Parameters:

Name Type Description Default
X

Features, shape (n, d) (or (n,)).

required
y

Integer class labels, shape (n,).

required
max_depth int

Maximum depth of the fitted tree.

2
criterion str

"gini" or "entropy".

'gini'

entropy_impurity(y)

Shannon entropy −Σ pᵢ log₂(pᵢ) of the labels in y.

gini_impurity(y)

Gini impurity 1 − Σ pᵢ² of the labels in y.

kmeans_trace(X, k=2, max_iters=10, init=None)

Build the full Lloyd's-algorithm trace, clustering X into k groups.

Parameters:

Name Type Description Default
X

Data points, shape (n, d) (or (n,) for 1-D data).

required
k int

Number of clusters.

2
max_iters int

Safety cap on assign/update rounds.

10
init ndarray | None

Optional starting centroids, shape (k, d). Defaults to the first k points, which keeps the demo deterministic.

None

knn_trace(X_train, y_train, x_query, k=3)

Build the full distance/vote trace classifying one query point.

Parameters:

Name Type Description Default
X_train

Training features, shape (n, d) (or (n,)).

required
y_train

Training labels, shape (n,).

required
x_query

A single query point, shape (d,) (or a scalar for 1-D).

required
k int

Number of neighbors to vote.

3

linear_regression_trace(X, y)

Build the full normal-equation trace fitting ŷ = Xθ to (X, y).

logistic_regression_trace(X, y, lr=0.5, steps=3, theta0=None)

Build a trace of a few gradient-descent steps fitting logistic regression.

Parameters:

Name Type Description Default
X

Feature matrix, shape (n, d) (or (n,) for one feature).

required
y

Binary labels in {0, 1}, shape (n,).

required
lr float

Learning rate for each gradient step.

0.5
steps int

Number of gradient-descent steps to trace (kept small — a real fit would run until the loss plateaus).

3
theta0 ndarray | None

Optional starting weights (zeros by default).

None

accuracy(y_true, y_pred)

Fraction of predictions that exactly match the true labels.

accuracy_trace(y_true, y_pred)

Build the trace of accuracy = fraction of exactly-correct predictions.

confusion_matrix(y_true, y_pred, labels=None)

Confusion matrix (rows = true label, columns = predicted label).

confusion_matrix_trace(y_true, y_pred, labels=None)

Build the trace of the confusion matrix (rows = true, columns = predicted).

mse(y_true, y_pred)

Mean squared error between y_true and y_pred.

mse_trace(y_true, y_pred)

Build the trace of mean squared error.

precision_recall_f1(y_true, y_pred, positive_label=1)

Return {"precision": ..., "recall": ..., "f1": ...} for positive_label.

precision_recall_f1_trace(y_true, y_pred, positive_label=1)

Build the trace of binary precision, recall, and F1 for positive_label.

r2_score(y_true, y_pred)

R² (coefficient of determination) between y_true and y_pred.

r2_score_trace(y_true, y_pred)

Build the trace of the R² (coefficient of determination) score.

roc_auc(y_true, y_scores)

ROC-AUC for binary labels y_true given continuous y_scores.

roc_auc_trace(y_true, y_scores)

Build the trace of ROC-AUC via the Mann-Whitney rank-sum formula.

y_true must be binary (0/1); y_scores are the model's continuous scores or probabilities for the positive class.

naive_bayes_trace(X_train, y_train, x_query)

Build the full prior/likelihood/posterior trace classifying one query point.

Parameters:

Name Type Description Default
X_train

Training features, shape (n, d) (or (n,)).

required
y_train

Training labels, shape (n,).

required
x_query

A single query point, shape (d,) (or a scalar for 1-D).

required

pca_trace(X, n_components=1)

Build the full center/covariance/eigendecomposition/project trace.

Parameters:

Name Type Description Default
X

Data, shape (n, d).

required
n_components int

How many top principal components to keep (k <= d).

1

optimumai.search

Classical AI search — how agents find paths and choose moves.

Two families, four algorithms:

  • Path-finding (:mod:optimumai.search.uninformed, :mod:optimumai.search.informed) — :func:bfs, :func:dfs, :func:uniform_cost_search, :func:greedy_best_first, and :func:astar all search a :class:~optimumai.search.problem.Graph or :class:~optimumai.search.problem.GridWorld for a route from a start state to a goal state.
  • Adversarial search (:mod:optimumai.search.adversarial) — :func:minimax and :func:alpha_beta choose a move in a two-player game tree, with alpha-beta pruning finding the identical value while visiting fewer nodes.

Every function has a *_trace counterpart (e.g. :func:astar_trace) that returns a full :class:~optimumai.core.trace.Trace of the search: frontier contents, expansion order, and (for informed/adversarial search) the g/h/f values or the alpha-beta window at every step.

GameNode dataclass

A node in a small, explicit game tree.

A leaf has value set and no children; an internal node has children and no value. player records whose turn it is to move at this node ("max" or "min"), used only for the trace narration — the recursion itself alternates automatically.

Attributes:

Name Type Description
name str

Short label for tracing, e.g. "A" or "root".

value float | None

The static evaluation, only set on leaves.

children list[GameNode]

Child nodes reachable by one move, only set on internal nodes.

player Literal['max', 'min']

Whose turn it is to move at this node.

is_leaf property

True if this node has no children (a terminal position).

Graph dataclass

An explicit directed, weighted graph used as a search problem.

The adjacency structure is a plain dict[state, dict[neighbor, cost]] — no external graph library needed. Edges are directed: add both a -> b and b -> a for an undirected edge (see :meth:add_edge).

Attributes:

Name Type Description
adjacency dict[State, dict[State, float]]

{state: {neighbor: edge_cost, ...}, ...}.

heuristics dict[State, float]

Optional straight-line-style estimates {state: h(state)} for use with greedy best-first / A. Defaults to 0 everywhere (which makes A degrade gracefully to uniform-cost search).

Example

g = Graph() g.add_edge("A", "B", 1) g.add_edge("B", "C", 2) g.neighbors("A") {'B': 1} g.cost("B", "C") 2

add_edge(a, b, cost=1.0, bidirectional=True)

Add an edge a -> b with the given cost (and the reverse, by default).

neighbors(state)

Return {neighbor: edge_cost} reachable in one step from state.

cost(a, b)

Return the edge cost of the direct move a -> b.

heuristic(state, goal)

Return an estimate of the remaining cost from state to goal.

Falls back to 0 for states with no registered heuristic, which is always admissible (never overestimates) but uninformative.

states()

All states that appear as either an edge source or destination.

GridWorld dataclass

A 2-D grid path-finding problem: 4-connected moves, optional walls.

Coordinates are (row, col) with row increasing downward, matching how the grid is usually printed. Movement is restricted to the 4 orthogonal neighbors (no diagonals), each costing 1 by default, so BFS already finds shortest hop-count paths here — the more interesting question is how much less work A* does than BFS/UCS once a heuristic is added.

Attributes:

Name Type Description
width int

Number of columns.

height int

Number of rows.

walls set[Coord]

Set of blocked (row, col) cells that cannot be entered.

Example

gw = GridWorld(width=3, height=3, walls={(1, 1)}) sorted(gw.neighbors((0, 1))) [(0, 0), (0, 2), (1, 1)]

in_bounds(state)

True if state lies within the grid's rectangle.

is_free(state)

True if state is in bounds and not a wall.

neighbors(state)

Return {neighbor: 1.0} for the in-bounds, non-wall 4-neighbors.

Note the returned dict includes cells even when they happen to be walls filtered out — walls are excluded entirely, never returned.

cost(a, b)

Cost of moving from a to an adjacent free cell b (always 1).

heuristic(state, goal, kind='manhattan')

Estimate the remaining cost from state to goal.

manhattan (|dr| + |dc|) is admissible here because every move costs exactly 1 and can only change row or column by 1 — the true remaining cost can never be less than the Manhattan distance. euclidean is also admissible (straight-line distance is never longer than a path constrained to a grid) but less tight, so it prunes less and A* typically expands more nodes with it than with Manhattan on a 4-connected grid.

render(path=None)

Render the grid as text: # walls, * path, . free cells.

alpha_beta(node, maximizing=True, alpha=float('-inf'), beta=float('inf'))

Return the minimax value of node's game tree, computed with pruning.

alpha_beta_trace(node, maximizing=True, alpha=float('-inf'), beta=float('inf'))

Build the full trace of minimax with alpha-beta pruning.

Returns the identical root value :func:minimax_trace would (see the module docstring for why), while typically visiting far fewer nodes — the trace records every prune with the (alpha, beta) window active at the time.

minimax(node, maximizing=True)

Return the minimax value of the root of node's game tree.

minimax_trace(node, maximizing=True)

Build the full trace of plain minimax evaluation over a game tree.

astar(problem, start, goal)

Return the optimal path from start to goal via A* (admissible h assumed).

astar_trace(problem, start, goal)

Build the full trace of A* search (order by f(n) = g(n) + h(n)).

Optimal whenever problem.heuristic is admissible (never overestimates the true remaining cost) — see the module docstring for the proof sketch.

greedy_best_first(problem, start, goal)

Return a path from start to goal via greedy best-first search.

greedy_best_first_trace(problem, start, goal)

Build the full trace of greedy best-first search (order by h(n) alone).

bfs(problem, start, goal)

Return the fewest-edges path from start to goal (BFS).

bfs_trace(problem, start, goal)

Build the full trace of breadth-first search from start to goal.

dfs(problem, start, goal)

Return a path from start to goal found via DFS (not necessarily optimal).

dfs_trace(problem, start, goal)

Build the full trace of depth-first search from start to goal.

Return the minimum-cost path from start to goal (UCS / Dijkstra).

uniform_cost_search_trace(problem, start, goal)

Build the full trace of uniform-cost search (Dijkstra to a single goal).

optimumai.rl

Reinforcement learning — how an agent learns to act from reward alone.

Four building blocks, each handed off to the next:

  • :mod:optimumai.rl.mdp — the formal object (states, actions, transitions, rewards, discount) and the Bellman equation that "solves" it exactly when the environment model is known (value iteration, policy iteration).
  • :mod:optimumai.rl.q_learning — model-free temporal-difference learning (Q-learning, SARSA) when the model is not known, only lived experience.
  • :mod:optimumai.rl.policy_gradient — REINFORCE, which parameterizes the policy directly and differentiates through sampled actions via the score-function trick.
  • :mod:optimumai.rl.ppo — PPO's clipped surrogate objective, the stabilized policy-gradient update that powers the reinforcement-learning stage of RLHF (contrasted with the closed-form DPO loss in :mod:optimumai.frontier.rlhf).

MDP dataclass

A finite, discounted Markov Decision Process.

Attributes:

Name Type Description
states list[str]

State labels, e.g. ["s0", "s1", "s2"].

actions list[str]

Action labels, e.g. ["left", "right"].

transition ndarray

P[s_idx][a_idx] is a probability vector over next-state indices — transition[s][a][s'] = P(s'|s, a). Must sum to 1 per (s, a).

reward ndarray

R[s_idx][a_idx][s'_idx] = immediate reward for that transition.

gamma float

Discount factor, 0 <= gamma < 1.

terminal frozenset[int]

Optional set of state indices with no outgoing value (their value is fixed at 0 and they are excluded from the max/backup).

expected_backup(values, s, a)

The Bellman backup Σₛ' P(s'|s,a)[R(s,a,s') + γV(s')] for one (s, a).

BanditEnv dataclass

A stateless k-armed bandit: pick an arm, get a fixed reward.

Deterministic rewards keep the lesson focused on how REINFORCE shifts probability mass toward the best action, without the extra variance a stochastic reward would add on top of the policy's own sampling noise.

PPOSample dataclass

One (state, action) observation replayed from a rollout batch.

Attributes:

Name Type Description
old_logprob float

logπ_θ_old(a|s) — log-prob under the policy that generated this action (frozen for the whole batch).

new_logprob float

logπ_θ(a|s) — log-prob under the policy currently being optimized (changes across optimizer steps/epochs).

advantage float

Aₜ, how much better this action was than the state's baseline (e.g. from a critic / GAE — not recomputed here).

GridWorld dataclass

A tiny, fully deterministic grid: step off the edge and you don't move.

The agent starts at start and the episode ends on reaching goal (reward +goal_reward) or a cell in traps (reward trap_reward, also terminal). Every other step costs step_reward (encourages short paths). Deterministic transitions isolate the TD-learning behavior being taught here from the extra variance a stochastic MDP would add.

step(pos, action)

Apply action at pos; return (next_pos, reward, done).

policy_iteration(mdp, theta=1e-06, max_iterations=1000, explain=False, level=ExplainLevel.ENGINEER)

Return V^π* for mdp. explain=True prints the full trace.

policy_iteration_trace(mdp, theta=1e-06, max_iterations=1000)

Build the full trace of policy iteration converging to π* on mdp.

Alternates policy evaluation — solve the linear system V^π = R + γPV^π exactly for the current (possibly suboptimal) policy — with policy improvement — make the policy greedy w.r.t. that V^π — until the policy stops changing. Guaranteed to converge in a finite number of iterations for a finite MDP.

value_iteration(mdp, theta=1e-06, max_iterations=1000, explain=False, level=ExplainLevel.ENGINEER)

Return V* for mdp. explain=True prints the full trace.

value_iteration_trace(mdp, theta=1e-06, max_iterations=1000)

Build the full trace of value iteration converging to V* on mdp.

Repeatedly applies the Bellman optimality backup to every state, tracking the largest change Δ each sweep, and stops once Δ < theta. The greedy policy w.r.t. the converged values is then extracted.

reinforce(env=None, episodes=200, alpha=0.1, gamma=0.99, seed=0, explain=False, level=ExplainLevel.ENGINEER)

Return the learned action-probability vector. explain=True prints the trace.

reinforce_trace(env=None, episodes=200, alpha=0.1, gamma=0.99, seed=0)

Build the full trace of REINFORCE learning a softmax policy on env.

Each "episode" here is a single-step trajectory (state, action, reward) — enough to show the sample → return → score-function-gradient → update loop without the bookkeeping of multi-step return-to-go, which :func:optimumai.rl.ppo.ppo_trace picks up with real advantages instead.

ppo_clip(batch=None, epsilon=0.2, explain=False, level=ExplainLevel.ENGINEER)

Return the PPO clipped loss for batch. explain=True prints the trace.

ppo_clip_trace(batch=None, epsilon=0.2)

Build the full trace of the PPO clipped surrogate objective on batch.

Walks every sample through the ratio, the unclipped and clipped terms, whether/why clipping engaged, and the final batch-averaged loss (L = -mean(L^CLIP) since optimizers minimize).

q_learning_trace(world=None, episodes=500, alpha=0.5, gamma=0.9, epsilon=0.1, seed=0)

Build the full trace of tabular Q-learning on world (default: a 3x3 grid).

sarsa(world=None, episodes=500, alpha=0.5, gamma=0.9, epsilon=0.1, seed=0, explain=False, level=ExplainLevel.ENGINEER)

Return the learned Q-table. explain=True prints the full trace.

sarsa_trace(world=None, episodes=500, alpha=0.5, gamma=0.9, epsilon=0.1, seed=0)

Build the full trace of tabular SARSA on world (default: a 3x3 grid).

optimumai.nlp

Classical NLP — the statistical machinery beneath modern language models.

Before attention, before embeddings even, NLP was tokenization, counting, and dynamic programming. This package covers the fundamentals that still power production search/retrieval systems and explain why the neural approach in :mod:optimumai.transformers and :mod:optimumai.embeddings works:

  • :mod:optimumai.nlp.bpe — byte-pair encoding, how text becomes tokens
  • :mod:optimumai.nlp.tfidf — TF-IDF, weighting words by how distinguishing they are
  • :mod:optimumai.nlp.ngram — n-gram language models, counting your way to "next word"
  • :mod:optimumai.nlp.edit_distance — Levenshtein distance via dynamic programming
  • :mod:optimumai.nlp.word2vec — skip-gram, the seed idea behind every learned embedding

BPETokenizer

Learn BPE merges from a corpus, then encode new words with them.

Parameters:

Name Type Description Default
num_merges int

How many merge rules to learn during :meth:train.

10

train(corpus)

Learn up to num_merges merge rules from a list of training words.

encode(word)

Apply learned merges (in training order) to tokenize word.

NGramModel

A counting-based n-gram language model with add-k smoothing.

Parameters:

Name Type Description Default
n int

Order of the model (2 = bigram, 3 = trigram, ...). Must be >= 1.

2
k float

Add-k smoothing constant (k=1 is classic Laplace smoothing).

1.0

vocab_size property

Number of distinct tokens seen during :meth:fit (including BOS/EOS).

fit(corpus)

Count n-grams (and their contexts) over a list of sentences.

prob(context, word)

Add-k smoothed P(word | context) for a (n-1)-length context.

sentence_log_prob(sentence)

Sum of log P(w_i | context) over every n-gram in a padded sentence.

perplexity(corpus)

PPL = exp(-(1/N) * sum of log-probs) over every n-gram in corpus.

TfidfVectorizer

Fit a vocabulary + idf table on a corpus, then transform documents to vectors.

Mirrors the fit/transform shape of familiar vectorizer APIs, but stays numpy + stdlib only, matching the "tf * smoothed-idf" formula documented in :func:tfidf_trace.

fit(docs)

Learn the vocabulary and idf weights from docs.

transform(docs)

Map each document in docs to a tf-idf row using the fitted idf.

fit_transform(docs)

Fit on docs then transform them in one call.

SkipGramModel

A tiny skip-gram word2vec model trained with full-softmax SGD.

Parameters:

Name Type Description Default
vocab list[str]

The (deduplicated) vocabulary; row i of both embedding matrices belongs to vocab[i].

required
dim int

Embedding dimensionality.

4
seed int

Seed for the initial embedding matrices.

0

predict(center)

Return P(context | center) over the whole vocabulary.

step(center, context, lr=0.1)

One full-softmax SGD step on a single (center, context) pair.

Returns the cross-entropy loss before the update.

most_similar(word, k=3)

Rank the vocabulary by cosine similarity to word in w_in space.

bpe_trace(corpus, num_merges, encode_word)

Train BPE on corpus and trace each merge round, then encode encode_word.

edit_distance_trace(a, b)

Build the full DP-table + backtrace trace for the Levenshtein distance.

edit_script(a, b)

Return the backtraced edit script turning a into b.

Each element is (op, from_char, to_char) with op in {"match", "substitute", "delete", "insert"}; "-" marks the absent side of an insert/delete.

ngram_trace(train_corpus, test_sentence, n=2, k=1.0)

Fit an n-gram model on train_corpus and trace its PPL on test_sentence.

perplexity(model, corpus)

Convenience wrapper: model.perplexity(corpus).

See :func:optimumai.evaluation.perplexity.perplexity for the general formula applied to a bare list of per-token probabilities; this variant is specific to a fitted :class:NGramModel scoring a corpus of sentences.

tfidf_trace(docs)

Build the full TF, IDF, and weighted-matrix trace for a document set.

skipgram_pairs(corpus, window=1)

Build every (center, context) pair within window of each center word.

word2vec_trace(corpus, window=1, dim=4, steps=20, lr=0.1, seed=0)

Build (center, context) pairs, train a few SGD steps, and trace the whole thing.

optimumai.vision

Computer vision fundamentals — how a CNN sees an image.

2D convolution (the local-pattern detector), pooling (translation-tolerant downsampling), Sobel edge detection (a fixed convolution finding boundaries), and a tiny CNN forward pass (conv -> ReLU -> pool -> flatten -> dense -> softmax) that ties the others together and headlines the tensor-shape story.

cnn_forward(image, kernel, dense_weights, dense_bias, pool_size=2)

Run one image through conv -> ReLU -> max-pool -> flatten -> dense -> softmax.

cnn_forward_trace(image, kernel, dense_weights, dense_bias, pool_size=2)

Build the full trace of a tiny CNN forward pass, headlined by shape flow.

dense(flat, weights, bias)

A fully-connected layer: logits = weights @ flat + bias.

relu(z)

Elementwise max(0, z) — zero out negative activations, keep the rest.

conv2d(x, w, stride=1, padding=0, mode='cross-correlate')

Slide kernel w over image x and return the output feature map.

mode="cross-correlate" (default) is what every deep learning framework calls "convolution": no kernel flip. mode="convolve" flips the kernel 180° first, matching the classical signal-processing definition.

conv2d_trace(x, w, stride=1, padding=0, mode='cross-correlate')

Build the full trace of sliding w over x, window by window.

sobel_edges(x, padding=1)

Return (magnitude, orientation) gradient maps for image x.

orientation is in radians, from atan2(Gy, Gx) (range (-π, π]).

sobel_edges_trace(x, padding=1)

Build the full trace of Sobel edge detection: Gx, Gy, magnitude, orientation.

avg_pool2d(x, kernel_size=2, stride=None)

Downsample x by averaging each kernel_size×kernel_size window.

max_pool2d(x, kernel_size=2, stride=None)

Downsample x by taking the max of each kernel_size×kernel_size window.

pool2d_trace(x, kernel_size=2, stride=None, how='max')

Build the full trace of pooling x, window by window.

optimumai.evaluation

LLM evaluation — how do you score a model's output when there's no single right answer?

Four complementary angles: surface-overlap text metrics (BLEU/ROUGE/F1/EM), intrinsic language-model quality (perplexity), whether a model's stated confidence can be trusted (calibration/ECE), and whether a generated answer stays grounded in its source (a hallucination/faithfulness heuristic — see :mod:optimumai.evaluation.hallucination for why that last one is an educational proxy, not a solved problem).

ece(confidences, correct, n_bins=5)

Return the Expected Calibration Error in [0, 1] (lower is better calibrated).

ece_trace(confidences, correct, n_bins=5)

Build the full ECE trace: bin the predictions, then show each bin's confidence-vs-accuracy gap.

Parameters:

Name Type Description Default
confidences Sequence[float]

The model's stated confidence for each prediction, in [0, 1].

required
correct Sequence[bool]

Whether each prediction was actually correct.

required
n_bins int

Number of equal-width reliability bins (default 5, i.e. 20%-wide bins).

5

faithfulness_score(answer, context)

Return the claim-overlap faithfulness heuristic in [0, 1] (see module docstring).

faithfulness_trace(answer, context)

Build the full trace: which answer tokens are supported by the context, and the score.

Parameters:

Name Type Description Default
answer str

The model's generated answer to score.

required
context str

The source text the answer was supposed to be grounded in (e.g. retrieved documents, a system prompt's reference material).

required

unsupported_spans(answer, context)

Return the answer's content tokens absent from the context's vocabulary.

These are the heuristic's flagged candidates for hallucinated content — treat them as "worth checking," not as confirmed fabrications (see module docstring).

perplexity_trace(probs)

Build the full trace: per-token surprisal → cross-entropy → perplexity.

Parameters:

Name Type Description Default
probs Iterable[float]

The probability the model assigned to the actual next token at each position (i.e. p(tokenᵢ | tokens_{<i}) for the ground truth sequence), one value per token. Must be in (0, 1].

required

bleu(candidate, reference, max_n=4)

Return the BLEU score of candidate against reference in [0, 1].

bleu_trace(candidate, reference, max_n=4)

Build the full BLEU trace: per-order clipped precision, then the brevity penalty.

exact_match(candidate, reference)

1.0 if the normalized token sequences are identical, else 0.0 (SQuAD-style).

rouge_l(candidate, reference)

Return the ROUGE-L F-measure in [0, 1].

rouge_l_trace(candidate, reference)

Build the ROUGE-L trace: LCS length → recall/precision/F-measure.

rouge_n(candidate, reference, n=1)

Return ROUGE-N recall (the conventionally reported value) in [0, 1].

rouge_n_trace(candidate, reference, n=1)

Build the ROUGE-N trace: recall/precision/F1 of order-n n-gram overlap.

token_f1(candidate, reference)

Return the token-level F1 of candidate against reference in [0, 1].

token_f1_trace(candidate, reference)

Build the token-F1 trace: bag-of-tokens precision/recall/F1 (SQuAD-style).

optimumai.prompting

Prompt-engineering patterns — constructed and explained offline.

Every pattern in this package builds a prompt deterministically, step by step, with no live LLM calls: the point is to see exactly how the prompt string is assembled and why the pattern helps (or where it fails), not to call out to a model. Each submodule exposes a <name>_trace function (returns a :class:~optimumai.core.trace.Trace), a thin <name> wrapper (returns the assembled prompt string, or renders the trace if explain=True), and a demo function for the curriculum/CLI.

chain_of_thought_trace(task, examples=None, trigger=DEFAULT_TRIGGER)

Build a CoT prompt: optional worked exemplars, task, reasoning scaffold, answer slot.

few_shot_trace(task, examples, instruction='')

Build a few-shot prompt: instruction + K exemplars + query, one exemplar at a time.

react_trace(task, tool_name='search', tool_description='search[query] — looks up query and returns a short snippet.')

Build a ReAct prompt: tool spec, task, and one Thought/Action/Observation cycle.

self_consistency_trace(task, sampled_answers=None)

Build the shared CoT prompt, then tally N deterministic toy sampled answers.

structured_output_trace(task, schema, example_completion=None)

Build a schema-constrained prompt and validate a toy completion against it.

zero_shot_trace(task, instruction='Complete the task below.', role=_DEFAULT_ROLE)

Build a zero-shot prompt: role + instruction + task, step by step.

optimumai.augmented_rnns

Attention and Augmented Recurrent Neural Networks.

Based on distill.pub's Attention and Augmented Recurrent Neural Networks (2016): three RNN-era ideas that gave networks capabilities beyond a fixed hidden state — attention as differentiable memory access, Neural Turing Machines' external read/write memory, and Adaptive Computation Time's learned, variable compute. All three converge on the same core trick (score, softmax, weighted blend) that later became transformer attention; see :mod:optimumai.transformers.attention for that descendant.

NTMMemory

A tiny Neural Turing Machine memory bank with content-based addressing.

Wraps :func:ntm_read / :func:ntm_write as stateful operations over a fixed-size memory matrix, mirroring how a real NTM controller would issue a read followed by a write each timestep.

read(key)

Content-addressed read from the current memory state.

write(key, erase, add)

Content-addressed write; updates and returns the new memory state.

adaptive_computation_time(halting_logits, eps=0.01)

Run ACT's halting mechanism over a sequence of ponder-step logits.

Parameters:

Name Type Description Default
halting_logits Iterable[float]

Raw (pre-sigmoid) halting score at each ponder step, in order. Must be non-empty.

required
eps float

Small tolerance; pondering stops once the cumulative halting probability reaches 1 - eps (or the logits run out).

0.01

Returns:

Type Description
dict[str, float | int | ndarray]

A dict with halting_probs (sigmoid of each logit up to and

dict[str, float | int | ndarray]

including the halt step), cumulative (running sum at each step),

dict[str, float | int | ndarray]

halt_step (1-based index of the step that triggered halting),

dict[str, float | int | ndarray]

remainder (weight given to the final step), and ponder_cost

dict[str, float | int | ndarray]

(halt_step + remainder).

adaptive_computation_time_trace(halting_logits, eps=0.01)

Build the full trace of ACT's ponder-and-halt mechanism.

Shows the per-step halting probabilities, the cumulative sum, the step at which pondering halts, the leftover remainder, and the resulting ponder cost (the term added to the training loss).

attention_read(query, memory)

Read a weighted blend of memory rows using content-based attention.

Parameters:

Name Type Description Default
query ndarray

A 1-D vector of shape (d,) — "what am I looking for?".

required
memory ndarray

A 2-D array of shape (n, d) — n memory slots of dimension d — "what is stored?".

required

Returns:

Type Description
ndarray

The blended read vector, shape (d,).

attention_read_trace(query, memory)

Build the full trace of a content-based memory read.

Shows the relevance scores, the softmax attention weights (which sum to 1), and the resulting blended read vector.

ntm_read(key, memory, beta=1.0)

Content-addressed read: blend memory rows by cosine similarity to key.

Parameters:

Name Type Description Default
key ndarray

1-D vector of shape (d,) used to address memory.

required
memory ndarray

2-D array of shape (n, d) — the memory bank.

required
beta float

Sharpness ("key strength"); larger values make addressing more peaked around the best-matching slot.

1.0

Returns:

Type Description
ndarray

The read vector, shape (d,).

ntm_trace(memory, read_key, write_key, erase, add, beta=1.0)

Build the full trace of an NTM read followed by a write.

Shows the content-based addressing weights for both operations, the read vector, and the memory bank before/after the write.

ntm_write(memory, key, erase, add, beta=1.0)

Content-addressed write: erase then add, blended by addressing weights.

Mᵢ ← Mᵢ · (1 − wᵢ · erase) + wᵢ · add for every slot i.

Parameters:

Name Type Description Default
memory ndarray

2-D array of shape (n, d) — the memory bank before the write.

required
key ndarray

1-D vector of shape (d,) used to address memory.

required
erase ndarray

1-D vector of shape (d,) in [0, 1] — what to clear.

required
add ndarray

1-D vector of shape (d,) — what to write in.

required
beta float

Addressing sharpness, as in :func:ntm_read.

1.0

Returns:

Type Description
ndarray

The new memory bank, shape (n, d).

optimumai.curriculum

The OptimumAI course — a first-principles AI learning path.

Course

Navigable view over the ordered curriculum.

tracks()

Lessons grouped by track, preserving order.

next_incomplete(completed)

The first lesson (in course order) not yet in completed.

Lesson dataclass

A single, runnable step in the course.

run(level='intermediate')

Render this lesson's trace at level and return it.

optimumai.quiz

Quiz / active-recall engine — test yourself, don't just reread.

Question dataclass

A single multiple-choice question tied to one curriculum lesson.

answer is the index into choices of the correct option, and explanation is a one-sentence justification shown after answering.

Quiz

A gradeable quiz over the questions for a single lesson.

questions property

The questions for this lesson, in bank order.

grade(answers)

Grade answers against the correct indices.

A length mismatch is tolerated: any question without a corresponding answer (and any answer past the last question) is scored as wrong.

check(index, choice)

Return whether choice is the correct option for question index.

QuizResult dataclass

The outcome of grading one quiz attempt.

per_question records, in order, whether each question was answered correctly, so a caller can highlight exactly what to revisit.

score property

Fraction correct in [0.0, 1.0] (0.0 for an empty quiz).

all_questions()

A flat list of every question across all lessons (for spaced repetition).

available_quizzes()

Sorted lesson ids that currently have at least one question in the bank.

optimumai.review

Spaced-repetition review (SM-2) over the OptimumAI course.

ReviewScheduler

Schedule and surface spaced-repetition reviews over a ProgressTracker.

record(lesson_id, quality, now=None)

Grade a review (0–5), persist the new schedule, and return it.

due(candidates=None, now=None)

Lessons whose review is due (never-reviewed lessons are due immediately).

ReviewState dataclass

SM-2 state for one lesson.

sm2(state, quality, now=None)

Apply one SM-2 update given a recall quality in 0–5.

optimumai.llm

Real token generation via local Ollama, Hugging Face, Anthropic, or a toy model.

available_providers()

Which generation providers are usable right now (toy is always last).

generate_trace(prompt, provider='auto', model=None, max_tokens=64, temperature=0.7)

Generate and return a :class:Trace of input → output tokens.

optimumai.visualization

Visualization: terminal traces, plus matplotlib graphs (the [viz] extra).

The plotting functions import matplotlib lazily, so importing this package never requires matplotlib — you only need it when you actually draw.

landscape_demo(out=None)

The headline demo: gradient descent tumbling down the Rosenbrock banana.

plot_loss_landscape(func='bowl', out=None, start=(-1.8, 1.8), lr=0.02, steps=40, kind='both')

Draw a loss landscape (contour, surface, or both) with a descent path.

plot_activation(name='gelu', xlim=(-6.0, 6.0), out=None)

Plot an activation and its (numeric) derivative on a single axes.

plot_attention(text=None, out=None)

Heatmap of scaled dot-product self-attention weights with token labels.

plot_embeddings(words=None, dim=16, out=None)

Scatter seeded word embeddings after a numpy-SVD PCA down to 2-D.

plot_heatmap(matrix, title='', labels=None, out=None)

imshow a 2-D array with a colorbar; annotate cells when the matrix is small.

plot_softmax_temperature(logits=(2.0, 1.0, 0.1), temperatures=(0.5, 1.0, 2.0, 5.0), out=None)

Grouped bars of softmax(logits / T) — sharp at low T, flat at high T.

plot_training_curve(losses=None, out=None)

Plot training loss vs iteration on a log-y axis (runs the demo if needed).

render_trace(trace, level=ExplainLevel.INTERMEDIATE, console=None)

Pretty-print trace at the requested detail level.

optimumai.circuit

The circuit view — render a computation graph like an electrical circuit.

Build a :class:FlowGraph from a :class:~optimumai.autograd.value.Value or a user expression, then render it as interactive HTML, Graphviz DOT, or a terminal table — with each node's forward data and backward gradient on the wires.

FlowGraph dataclass

A flattened, renderable view of a :class:Value computation DAG.

Attributes:

Name Type Description
nodes list[FlowNode]

Every node — value nodes and operator pseudo-nodes.

edges list[tuple[str, str]]

Directed wires as (from_id, to_id) tuples.

value_nodes()

Return only the real value nodes (leaves + intermediates).

op_nodes()

Return only the operator pseudo-nodes.

FlowNode dataclass

A single node in a :class:FlowGraph.

Attributes:

Name Type Description
id str

Stable identity string (str(id(value)) for value nodes, "<value_id>_op" for operator pseudo-nodes).

label str

Human label — a variable name ("a"), the result name ("L"), or an operator symbol ("*").

data float

Forward value (the number that flows forward through the wire).

grad float

Backward gradient (∂L/∂node — flows backward).

op str

The op that produced this value ("" for leaves / op nodes carry the symbol here too).

is_op bool

True for operator pseudo-nodes, False for value nodes.

build_from_expression(expression, variables=None, run_backward=True)

Compile a safe arithmetic expression into a Value graph + FlowGraph.

Parameters:

Name Type Description Default
expression str

An arithmetic string using variable names, numbers, and the operators + - * / ** with parentheses, e.g. "(a*b + c) * f".

required
variables dict[str, float] | None

Optional map of name → value. Names not present default to 1.0. Each distinct name becomes a labelled leaf Value.

None
run_backward bool

If True (default), run .backward() on the result so the graph carries gradients.

True

Returns:

Type Description
Value

(result_value, flow_graph) — the labelled result Value (label

FlowGraph

"L") and its flattened :class:FlowGraph.

Raises:

Type Description
ValueError

If the expression is empty, cannot be parsed, or contains disallowed syntax (anything beyond names / numbers / arithmetic).

build_from_value(root)

Flatten a :class:Value DAG into a :class:FlowGraph.

Walks root._topo() (dependencies before dependents). Every Value becomes a value node. Every Value that was produced by an op (has both _op and children) additionally gets an operator pseudo-node wired in between it and its children — mirroring micrograd's draw_dot::

child ─┐
       ├─▶ (op) ─▶ result
child ─┘

Leaves (no _op / no children) get no op node.

to_dot(fg)

Emit a Graphviz DOT description of fg (left-to-right layout).

Value nodes are records showing label | data d.dd | grad g.gg; operator pseudo-nodes are small ellipses carrying the op symbol.

to_html(fg, out)

Write a self-contained interactive circuit page to out; return path.

Uses vis-network from a CDN (single <script> tag — no other local assets). Value nodes show label, data (blue) and grad (orange) on multiple lines; operator pseudo-nodes are small circles. Edges are the directed "wires". Physics is enabled for a self-organising layout.

to_terminal(fg, console=None)

Print fg as "the circuit" in the terminal using Rich.

Nodes are listed in topological order (as stored). Each value row shows its label, the producing op, its data (blue) and grad (orange), plus a one-line summary of the wires flowing into it.