API Reference¶
Auto-generated from the docstrings of the public optimumai package.
Top-level exports¶
Everything in optimumai.__all__ is the public API and follows
semantic versioning:
# Linear algebra
from optimumai import Vector, Matrix
# Probability & calculus
from optimumai import softmax, derivative, gradient, integrate
# Autograd
from optimumai import Value
# Neural networks
from optimumai import MLP, Adam, minimize
# Transformers
from optimumai import Attention, MultiHeadAttention, TransformerBlock, TextPipeline
# World models & interpretability
from optimumai import JEPA, superposition
# Embeddings, RAG, diffusion
from optimumai import embedding_lookup, nearest_neighbors, RAGPipeline, forward_diffusion
# Systems
from optimumai import kv_cache_size, vram_estimate
# Generation & tutor
from optimumai import generate
# Course & progress
from optimumai import COURSE, ProgressTracker, Quiz, ReviewScheduler
# Visualization
from optimumai import render_concept, editable_plot
# Kernels
from optimumai import KernelWorkbench, GpuSim
# Trace types
from optimumai import Trace, ExplainLevel, Lesson, Course, Workbook
Module reference¶
optimumai
¶
OptimumAI — unlock the math behind AI.
Every operation, from a dot product to a full transformer block, can be run with
explain=True to produce a step-by-step computation trace, a terminal
visualization, and the context for why AI uses it.
>>> from optimumai import Vector
>>> Vector([1, 2, 3]).dot(Vector([4, 5, 6]), explain=True)
32.0
v0.2 adds the fundamentals behind modern AI: a micrograd-style autograd engine, calculus, optimizers, neural networks with real backprop, multi-head attention, transformer blocks, LeCun's JEPA world model, and Anthropic-style superposition.
v0.3 turns it into a learning path: a structured Course, progress tracking, a Streamlit dashboard, plus embeddings, RAG, diffusion, and an optional LLM tutor.
v0.4 adds the foundations of the stack: tensors & integration, the PyTorch and JAX programming models, and the systems layer — the CUDA execution & memory model, tiled matmul kernels, the KV cache, and a VRAM budget calculator.
v0.5 makes it interactive: feed your own input and watch it flow — a REPL, a text-to-transformer pipeline, op comparisons, parameter sweeps, and symbolic differentiation of your own equations.
v0.6 adds graphs: matplotlib figures (via optimumai.visualization) for
activation curves, attention heatmaps, embedding scatter, and a 3D loss
landscape with the gradient-descent trajectory carved across it.
v0.7 adds the circuit: render any expression or Value graph as a computation
"circuit" (interactive HTML, Graphviz DOT, or terminal) via optimumai.circuit,
with data and gradients flowing the wires.
v0.8 adds frontier concepts: FlashAttention (IO-aware tiling + online softmax), quantization (int8/int4), LoRA (parameter-efficient fine-tuning), and DPO (preference alignment) — how today's large models are actually built and run.
v0.9 makes it a learning product, grounded in cognitive science: a quiz / active-recall engine (the testing effect), spaced-repetition review (SM-2), guided onboarding, and course search.
v0.10 is hands-on: GPU kernels from scratch on a pure-Python simulator (write & grade your own), an editable equation↔graph in the browser, animated GIF export, and compute-the-answer exercises. Get your hands dirty.
v1.0 is the stable release: real token generation (Ollama / Hugging Face / Anthropic / a toy fallback), interactive drag-the-inputs circuits, a visualize-any-concept registry (PNG + GIF), a notebook launcher, and a docs site.
v1.1 broadens from the deep-learning stack to the whole field: classical machine learning (regression, trees, k-means, KNN, naive Bayes, PCA, metrics), classical AI search (BFS/DFS/UCS, A*, minimax + alpha-beta), reinforcement learning (MDPs & the Bellman equation, Q-learning/SARSA, REINFORCE, PPO), NLP fundamentals (BPE, TF-IDF, n-grams, edit distance, word2vec), computer vision (convolution, pooling, Sobel edges, a tiny CNN), and LLM evaluation (BLEU/ROUGE, perplexity, calibration, and a candid hallucination/faithfulness heuristic) — each an explainable trace.
v1.2 makes it interactive & explained: prompt-engineering patterns (zero/few-shot,
chain-of-thought, ReAct, self-consistency, structured output), the augmented-RNN
lineage from distill.pub (attention as differentiable memory, Neural Turing
Machines, Adaptive Computation Time), self-contained interactive HTML playgrounds
(a Transformer-Explainer-style attention widget, plus k-means and A*), and a
per-concept plot/GIF gallery wired into the visualize registry.
v1.3 makes it flow: a Plot Studio (feed numbers, get any chart — bar,
histogram, scatter, box, line, pie, violin — plus the exact matplotlib + numpy
code on screen), and a flows subpackage of distill.pub-style interactive
circuit-flow diagrams (a transformer forward pass, scaled dot-product attention,
TF-IDF, and word2vec) rendered as self-contained, offline HTML.
v1.4 teaches the tools: runnable, explained tutorials for NumPy,
matplotlib, PyTorch (built on OptimumAI's own engine so it runs torch-free,
with the real torch code shown), and LLM fine-tuning (a numpy SFT → LoRA → DPO
toy pipeline plus the production HF/PEFT/TRL code). Walk any of them in the
terminal (optimumai tutorial numpy) or export to a notebook — and read the
matching in-depth guides on the docs site.
v1.5 hardens the interactive layer with OptiX, a typed, unit-tested TypeScript
widget kit (authored in web/, compiled by esbuild to a single self-contained
IIFE shipped in the wheel) — replacing hand-written inline JS. Its debut is a
TensorFlow-Playground-style neural-net playground (optimumai playground nn):
pick a 2-D dataset, train a tiny MLP, and watch the decision boundary form.
v1.6 adds the Concept Explorer: 30 foundational AI/ML concepts (attention,
backprop, gradients, Adam/AdamW, activations, layer norm, PCA, k-means, RL,
tokenization, transformer blocks, and more) each rendered as a DAG you step
through, with a KaTeX formula and a runnable optimumai code snippet for
every step (optimumai explain <concept>), plus a searchable landing page
listing all of them (optimumai explore).
Public-API stability¶
Everything exported from the top-level optimumai namespace (the names in
__all__ below) is considered stable and follows semantic versioning: no
breaking changes without a major-version bump. Submodule internals and any
name prefixed with _ may change between minor releases.
Matrix
¶
Vector
¶
A 1-D array with step-by-step explanations of its operations.
dot_trace(other)
¶
Build the full trace of self · other.
dot(other, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Inner product. Set explain=True to print the full trace.
norm_trace()
¶
Euclidean (L2) norm, ||a|| = √Σ aᵢ².
cosine_similarity_trace(other)
¶
Cosine similarity (a · b) / (||a|| ||b||) — the RAG workhorse.
NTMMemory
¶
A tiny Neural Turing Machine memory bank with content-based addressing.
Wraps :func:ntm_read / :func:ntm_write as stateful operations over a
fixed-size memory matrix, mirroring how a real NTM controller would issue
a read followed by a write each timestep.
Value
¶
A differentiable scalar node in an autograd graph.
zero_grad()
¶
Reset the gradient of every node reachable from this one.
backward()
¶
Run reverse-mode autodiff: seed grad=1 and apply the chain rule.
backward_trace()
¶
Like :meth:backward, but record how each gradient is produced.
For every node processed (output → inputs) we snapshot its children's
gradients before and after applying the local _backward, so the delta
we log is the exact contribution the chain rule pushed to each child.
backprop(explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Convenience: run backward, optionally rendering the chain-rule trace.
ExplainLevel
¶
Bases: str, Enum
How much detail a rendered trace should surface.
Step
dataclass
¶
One intermediate move in a computation.
Attributes:
| Name | Type | Description |
|---|---|---|
index |
int
|
1-based position in the trace. |
title |
str
|
Short label, e.g. |
expression |
str
|
Human-readable computation, e.g. |
value |
Any
|
The numeric result of this step (scalar / vector / matrix). |
detail |
str
|
Optional longer note shown at higher explain levels. |
Trace
dataclass
¶
A complete, replayable record of an operation.
Attributes:
| Name | Type | Description |
|---|---|---|
op |
str
|
Name of the operation, e.g. |
result |
Any
|
The final output value. |
steps |
list[Step]
|
Ordered intermediate steps. |
formula |
str
|
The closed-form formula, e.g. |
why_ai |
list[str]
|
Bullet points on where this shows up in real AI systems. |
complexity |
str
|
Big-O note, surfaced at engineer/researcher levels. |
meta |
dict[str, Any]
|
Free-form extra data (shapes, hyper-params, ...). |
add(title, expression, value=None, detail='')
¶
Append a step and return self for fluent chaining.
render(level=ExplainLevel.INTERMEDIATE, console=None)
¶
Pretty-print this trace to the terminal and return the result.
Imported lazily so the math layer never hard-depends on the UI layer.
last()
¶
Return the final recorded step (raises if the trace is empty).
Course
¶
Lesson
dataclass
¶
A single, runnable step in the course.
run(level='intermediate')
¶
Render this lesson's trace at level and return it.
Workbook
¶
Serve and grade compute-the-value exercises (optionally for one lesson).
grade(exercise_id, value)
¶
Grade a submitted value (absolute tolerance, relative fallback for large answers).
KernelWorkbench
¶
GpuSim
¶
A serial simulator of a SIMT GPU launch.
Instantiate once per kernel run (or reuse and call :meth:reset). The core is
:meth:launch, which iterates every (block, thread) in the requested grid,
hands each a :class:ThreadCtx, and tallies memory traffic into self.stats.
reset()
¶
Clear all instrumentation so the simulator can be launched again.
launch(grid, block, kernel, *, flops=0)
¶
Run kernel once per thread in the grid × block launch.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
grid
|
int | tuple[int, int]
|
Blocks in the grid — |
required |
block
|
int | tuple[int, int]
|
Threads per block — |
required |
kernel
|
Callable[[ThreadCtx], None]
|
A function taking a single :class: |
required |
flops
|
int
|
Total floating-point ops the kernel performs, recorded for arithmetic-intensity reporting. Purely informational. |
0
|
Returns:
| Type | Description |
|---|---|
MemoryStats
|
|
GpuSim
|
Kernels write their output into caller-owned arrays, so the "output |
tuple[MemoryStats, GpuSim]
|
buffer" is whatever array the caller passed in and the kernel stored to. |
Threads are visited block-by-block; within a block they are grouped into
warps of :data:WARP_SIZE, and coalescing is judged per warp (consecutive
lanes → consecutive addresses). 2-D blocks are flattened in row-major
(ty, tx) order for warp grouping, matching CUDA's lane numbering.
KNN
¶
k-nearest-neighbors classifier: store the data, vote at prediction time.
Example
model = KNN(k=1).fit([[0], [1], [10], [11]], [0, 0, 1, 1]) model.predict([[0.5], [10.5]]) array([0, 1])
PCA
¶
Principal Component Analysis via covariance eigendecomposition.
Example
model = PCA(n_components=1).fit([[0, 0], [1, 1], [2, 2], [3, 3]]) model.transform([[4, 4]]).shape (1, 1)
fit(X)
¶
Compute the mean and top principal components from X.
transform(X)
¶
Project new data onto the fitted principal components.
fit_transform(X)
¶
Fit on X and return the projected training data in one call.
explain(X, level=ExplainLevel.INTERMEDIATE)
¶
Fit, print the full center/covariance/eigen trace, and return the projection.
DecisionTree
¶
A shallow classification tree grown by greedy impurity-reduction splits.
Example
model = DecisionTree(max_depth=1).fit([[0], [1], [10], [11]], [0, 0, 1, 1]) model.predict([[0.5], [10.5]]) array([0, 1])
GaussianNB
¶
Gaussian Naive Bayes classifier.
Example
model = GaussianNB().fit([[0], [1], [10], [11]], [0, 0, 1, 1]) model.predict([[0.5], [10.5]]) array([0, 1])
KMeans
¶
Lloyd's-algorithm k-means clustering.
Example
model = KMeans(k=2).fit([[0], [1], [10], [11]]) model.predict([[0.5], [10.5]]) array([0, 1])
LinearRegression
¶
Ordinary least squares via the normal equation.
Example
model = LinearRegression().fit([[1], [2], [3]], [2, 4, 6]) model.predict([[4]]) array([8.])
LogisticRegression
¶
Binary logistic regression trained by batch gradient descent.
Example
model = LogisticRegression(lr=0.5, steps=200).fit([[0], [1], [2], [3]], [0, 0, 1, 1]) model.predict([[0.1], [2.9]]) array([0, 1])
fit(X, y)
¶
Run :attr:steps gradient-descent updates and store the final weights.
predict_proba(X)
¶
Return P(y=1|x) for new samples (must call :meth:fit first).
predict(X)
¶
Return the hard 0/1 label (threshold at 0.5).
explain(X, y, level=ExplainLevel.INTERMEDIATE)
¶
Fit, print the full trace (with :attr:steps gradient updates shown), return labels.
MLP
¶
A stack of dense layers. Hidden layers use activation; output is linear.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n_in
|
int
|
Number of input features. |
required |
n_outs
|
Sequence[int]
|
Sizes of each layer, e.g. |
required |
activation
|
str
|
Hidden-layer nonlinearity ( |
'tanh'
|
seed
|
int
|
Seed for reproducible weight initialization. |
0
|
BPETokenizer
¶
Learn BPE merges from a corpus, then encode new words with them.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
num_merges
|
int
|
How many merge rules to learn during :meth: |
10
|
NGramModel
¶
A counting-based n-gram language model with add-k smoothing.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n
|
int
|
Order of the model (2 = bigram, 3 = trigram, ...). Must be >= 1. |
2
|
k
|
float
|
Add-k smoothing constant ( |
1.0
|
vocab_size
property
¶
Number of distinct tokens seen during :meth:fit (including BOS/EOS).
fit(corpus)
¶
Count n-grams (and their contexts) over a list of sentences.
prob(context, word)
¶
Add-k smoothed P(word | context) for a (n-1)-length context.
sentence_log_prob(sentence)
¶
Sum of log P(w_i | context) over every n-gram in a padded sentence.
perplexity(corpus)
¶
PPL = exp(-(1/N) * sum of log-probs) over every n-gram in corpus.
SkipGramModel
¶
A tiny skip-gram word2vec model trained with full-softmax SGD.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
vocab
|
list[str]
|
The (deduplicated) vocabulary; row |
required |
dim
|
int
|
Embedding dimensionality. |
4
|
seed
|
int
|
Seed for the initial embedding matrices. |
0
|
predict(center)
¶
Return P(context | center) over the whole vocabulary.
step(center, context, lr=0.1)
¶
One full-softmax SGD step on a single (center, context) pair.
Returns the cross-entropy loss before the update.
most_similar(word, k=3)
¶
Rank the vocabulary by cosine similarity to word in w_in space.
TfidfVectorizer
¶
Fit a vocabulary + idf table on a corpus, then transform documents to vectors.
Mirrors the fit/transform shape of familiar vectorizer APIs, but stays
numpy + stdlib only, matching the "tf * smoothed-idf" formula documented
in :func:tfidf_trace.
SGD
¶
Vanilla stochastic gradient descent: w ← w − lr · ∂L/∂w.
Adam
¶
Adam: per-parameter adaptive steps from 1st/2nd gradient moments.
m ← β₁m + (1−β₁)g, v ← β₂v + (1−β₂)g² m̂ = m/(1−β₁ᵗ), v̂ = v/(1−β₂ᵗ) w ← w − lr · m̂ / (√v̂ + ε)
ProgressTracker
¶
Quiz
¶
A gradeable quiz over the questions for a single lesson.
questions
property
¶
The questions for this lesson, in bank order.
grade(answers)
¶
Grade answers against the correct indices.
A length mismatch is tolerated: any question without a corresponding answer (and any answer past the last question) is scored as wrong.
check(index, choice)
¶
Return whether choice is the correct option for question index.
RAGPipeline
¶
Bases: BaseOp
A minimal, deterministic retrieval-augmented-generation pipeline.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
corpus
|
list[str] | None
|
The documents to retrieve from. Defaults to five ML sentences. |
None
|
dim
|
int
|
Embedding dimension for the bag-of-hashed-words vectors. |
16
|
seed
|
int
|
Base seed; each word gets its own vector from |
0
|
ReviewScheduler
¶
MDP
dataclass
¶
A finite, discounted Markov Decision Process.
Attributes:
| Name | Type | Description |
|---|---|---|
states |
list[str]
|
State labels, e.g. |
actions |
list[str]
|
Action labels, e.g. |
transition |
ndarray
|
|
reward |
ndarray
|
|
gamma |
float
|
Discount factor, |
terminal |
frozenset[int]
|
Optional set of state indices with no outgoing value (their value is fixed at 0 and they are excluded from the max/backup). |
expected_backup(values, s, a)
¶
The Bellman backup Σₛ' P(s'|s,a)[R(s,a,s') + γV(s')] for one (s, a).
Attention
¶
Bases: BaseOp
Single-head scaled dot-product attention.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
d_k
|
int | None
|
Key/query dimension used for the |
None
|
TransformerBlock
¶
Bases: BaseOp
One pre-norm transformer decoder block (attention + feed-forward).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
d_model
|
int
|
Model / residual-stream dimension. |
required |
n_heads
|
int
|
Number of attention heads (must divide |
required |
d_ff
|
int | None
|
Hidden width of the feed-forward network (default |
None
|
MultiHeadAttention
¶
Bases: BaseOp
Multi-head scaled dot-product self-attention.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n_heads
|
int
|
Number of parallel attention heads. |
required |
d_model
|
int
|
Model dimension; must be divisible by |
required |
TextPipeline
¶
Bases: BaseOp
Push a user's own text through a full, tiny, untrained transformer.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
The prompt to run through the pipeline (must be non-empty). |
required |
layers
|
int
|
How many :class: |
2
|
d_model
|
int
|
Model / residual-stream dimension (must divide by |
16
|
n_heads
|
int
|
Number of attention heads per block. |
2
|
seed
|
int
|
Seed for the (fixed, untrained) embedding and output-head weights. |
0
|
tokenizer
|
str
|
|
'whitespace'
|
Tutor
¶
A friendly, optional LLM tutor for the math behind AI.
Safe to construct and call unconditionally: when no client library or key is available, methods return a clear message rather than raising.
available
property
¶
True when a client library and an API key are both available.
ask(question, system=None)
¶
Ask a free-form question. Never raises; returns guidance if unavailable.
explain(concept, level=ExplainLevel.INTERMEDIATE)
¶
Explain a math/AI concept at the requested detail level.
explain_trace(trace, level=ExplainLevel.INTERMEDIATE)
¶
Add LLM intuition on top of an OptimumAI :class:Trace.
JEPA
¶
Bases: BaseOp
A minimal Joint-Embedding Predictive Architecture.
Uses a fixed, seeded linear encoder f (shared by context and target) and
a linear predictor g living entirely in embedding space. Everything is
deterministic given seed so traces are reproducible.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_dim
|
int
|
Dimension of a raw input view (e.g. flattened patch). |
required |
embed_dim
|
int
|
Dimension of the abstract representation space. |
required |
seed
|
int
|
RNG seed for the encoder/predictor weights. |
0
|
forward(context, target, explain=False, level='engineer')
¶
Compute the energy. Set explain=True to print the four-stage trace.
demo(seed=0)
classmethod
¶
A tiny, reproducible example: two noisy views of one latent vector.
target = context + small noise mimics two augmentations of the same
image, so a well-behaved encoder keeps the energy small.
compare(op_a, op_b, input, explain=False, level=ExplainLevel.ENGINEER)
¶
Compare op_a and op_b on input. Set explain=True to print.
sweep(op, param, values, base_input=None, explain=False, level=ExplainLevel.ENGINEER)
¶
Sweep param of op across values. Set explain=True to print.
adaptive_computation_time(halting_logits, eps=0.01)
¶
Run ACT's halting mechanism over a sequence of ponder-step logits.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
halting_logits
|
Iterable[float]
|
Raw (pre-sigmoid) halting score at each ponder step, in order. Must be non-empty. |
required |
eps
|
float
|
Small tolerance; pondering stops once the cumulative halting
probability reaches |
0.01
|
Returns:
| Type | Description |
|---|---|
dict[str, float | int | ndarray]
|
A dict with |
dict[str, float | int | ndarray]
|
including the halt step), |
dict[str, float | int | ndarray]
|
|
dict[str, float | int | ndarray]
|
|
dict[str, float | int | ndarray]
|
( |
attention_read(query, memory)
¶
Read a weighted blend of memory rows using content-based attention.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
query
|
ndarray
|
A 1-D vector of shape |
required |
memory
|
ndarray
|
A 2-D array of shape |
required |
Returns:
| Type | Description |
|---|---|
ndarray
|
The blended read vector, shape |
ntm_read(key, memory, beta=1.0)
¶
Content-addressed read: blend memory rows by cosine similarity to key.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
key
|
ndarray
|
1-D vector of shape |
required |
memory
|
ndarray
|
2-D array of shape |
required |
beta
|
float
|
Sharpness ("key strength"); larger values make addressing more peaked around the best-matching slot. |
1.0
|
Returns:
| Type | Description |
|---|---|
ndarray
|
The read vector, shape |
ntm_write(memory, key, erase, add, beta=1.0)
¶
Content-addressed write: erase then add, blended by addressing weights.
Mᵢ ← Mᵢ · (1 − wᵢ · erase) + wᵢ · add for every slot i.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
memory
|
ndarray
|
2-D array of shape |
required |
key
|
ndarray
|
1-D vector of shape |
required |
erase
|
ndarray
|
1-D vector of shape |
required |
add
|
ndarray
|
1-D vector of shape |
required |
beta
|
float
|
Addressing sharpness, as in :func: |
1.0
|
Returns:
| Type | Description |
|---|---|
ndarray
|
The new memory bank, shape |
derivative(f, x, h=1e-05, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Numeric derivative of f at x (central difference).
gradient(f, point, h=1e-05, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Numeric gradient of f at point.
forward_diffusion(x0, timesteps=10, beta_start=0.0001, beta_end=0.2, seed=0, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Noise x0 to timestep T. Set explain=True to print the trace.
embedding_lookup(tokens, dim=4, seed=0, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Look up embedding rows for tokens. Set explain=True to print it.
nearest_neighbors(query, vocab, dim=4, seed=0, k=3, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Return the top-k cosine neighbours of query within vocab.
bleu(candidate, reference, max_n=4)
¶
Return the BLEU score of candidate against reference in [0, 1].
ece(confidences, correct, n_bins=5)
¶
Return the Expected Calibration Error in [0, 1] (lower is better calibrated).
exact_match(candidate, reference)
¶
1.0 if the normalized token sequences are identical, else 0.0 (SQuAD-style).
faithfulness_score(answer, context)
¶
Return the claim-overlap faithfulness heuristic in [0, 1] (see module docstring).
rouge_l(candidate, reference)
¶
Return the ROUGE-L F-measure in [0, 1].
rouge_n(candidate, reference, n=1)
¶
Return ROUGE-N recall (the conventionally reported value) in [0, 1].
token_f1(candidate, reference)
¶
Return the token-level F1 of candidate against reference in [0, 1].
kv_cache_size(n_layers, n_heads, head_dim, seq_len, batch=1, bytes_per_elem=2, kv_heads=None)
¶
Return the KV cache size in bytes (no trace).
kv_heads defaults to n_heads (Multi-Head Attention); pass a smaller
number for Grouped-Query Attention or 1 for Multi-Query Attention.
integrate(f, a, b, method='trapezoid', n=100, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Return ∫ f from a to b. Set explain=True to print the trace.
vram_estimate(params_billions, precision_bytes=2, training=True, optimizer='adam', batch=1, seq_len=2048, hidden=4096, n_layers=32, activation_factor=None)
¶
Return the estimated total VRAM in GB (no trace).
flash_attention(Q, K, V, block_size=2, explain=False, level=ExplainLevel.ENGINEER)
¶
Return flash attention output. Set explain=True to print the trace.
lora(d_in=8, d_out=8, rank=2, seed=0, explain=False, level=ExplainLevel.ENGINEER)
¶
Return the LoRA parameter reduction factor. explain=True prints the trace.
quantize(x, bits=8, scheme='symmetric', granularity='tensor')
¶
Quantize x to integer codes.
Returns:
| Type | Description |
|---|---|
|
|
|
|
|
dpo(prompt='Explain gravity.', beta=0.1, seed=0, explain=False, level=ExplainLevel.ENGINEER)
¶
Return the DPO loss for one preference triple. explain=True prints the trace.
run_repl()
¶
Start the interactive read-eval-print loop.
superposition(n_features, n_neurons, sparsity=0.7, seed=0, explain=False, level=ExplainLevel.ENGINEER)
¶
Run the toy superposition model, returning the recovered vector ĥ.
Set explain=True to print the encode/interference/decode trace.
generate(prompt, provider='auto', model=None, max_tokens=64, temperature=0.7)
¶
Generate a continuation of prompt and return the text.
provider="auto" picks the best available (Ollama → HF → Anthropic → toy).
minimize(loss_fn, params, optimizer, steps=50, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Minimize loss_fn over params and return the final parameter values.
softmax(x, temperature=1.0, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Return the softmax of x. Set explain=True to print the trace.
softmax_trace(x, temperature=1.0)
¶
Build the full trace of softmax(x) with an optional temperature.
policy_iteration(mdp, theta=1e-06, max_iterations=1000, explain=False, level=ExplainLevel.ENGINEER)
¶
Return V^π* for mdp. explain=True prints the full trace.
ppo_clip(batch=None, epsilon=0.2, explain=False, level=ExplainLevel.ENGINEER)
¶
Return the PPO clipped loss for batch. explain=True prints the trace.
reinforce(env=None, episodes=200, alpha=0.1, gamma=0.99, seed=0, explain=False, level=ExplainLevel.ENGINEER)
¶
Return the learned action-probability vector. explain=True prints the trace.
sarsa(world=None, episodes=500, alpha=0.5, gamma=0.9, epsilon=0.1, seed=0, explain=False, level=ExplainLevel.ENGINEER)
¶
Return the learned Q-table. explain=True prints the full trace.
value_iteration(mdp, theta=1e-06, max_iterations=1000, explain=False, level=ExplainLevel.ENGINEER)
¶
Return V* for mdp. explain=True prints the full trace.
alpha_beta(node, maximizing=True, alpha=float('-inf'), beta=float('inf'))
¶
Return the minimax value of node's game tree, computed with pruning.
astar(problem, start, goal)
¶
Return the optimal path from start to goal via A* (admissible h assumed).
bfs(problem, start, goal)
¶
Return the fewest-edges path from start to goal (BFS).
dfs(problem, start, goal)
¶
Return a path from start to goal found via DFS (not necessarily optimal).
greedy_best_first(problem, start, goal)
¶
Return a path from start to goal via greedy best-first search.
minimax(node, maximizing=True)
¶
Return the minimax value of the root of node's game tree.
uniform_cost_search(problem, start, goal)
¶
Return the minimum-cost path from start to goal (UCS / Dijkstra).
differentiate(expression, var='x', at=None, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Differentiate expression w.r.t. var. Set explain=True to print.
positional_encoding(seq_len, d_model, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Return the sinusoidal PE matrix. Set explain=True to print the trace.
get_tutorial(name)
¶
Load and build the tutorial called name (lazy import).
list_tutorials()
¶
The available tutorial names.
avg_pool2d(x, kernel_size=2, stride=None)
¶
Downsample x by averaging each kernel_size×kernel_size window.
cnn_forward(image, kernel, dense_weights, dense_bias, pool_size=2)
¶
Run one image through conv -> ReLU -> max-pool -> flatten -> dense -> softmax.
conv2d(x, w, stride=1, padding=0, mode='cross-correlate')
¶
Slide kernel w over image x and return the output feature map.
mode="cross-correlate" (default) is what every deep learning framework
calls "convolution": no kernel flip. mode="convolve" flips the kernel
180° first, matching the classical signal-processing definition.
max_pool2d(x, kernel_size=2, stride=None)
¶
Downsample x by taking the max of each kernel_size×kernel_size window.
sobel_edges(x, padding=1)
¶
Return (magnitude, orientation) gradient maps for image x.
orientation is in radians, from atan2(Gy, Gx) (range (-π, π]).
optix_js()
cached
¶
Return the OptiX bundle source (window.OptiX IIFE).
Raises FileNotFoundError if the compiled asset wasn't shipped — build it
with cd web && npm install && npm run build.
render_concept(name, fmt='png', out=None)
¶
Render name as fmt ("png" or "gif") to out (or a default path).
explain(concept, out=None, open_browser=True)
¶
Build an interactive DAG explainer (formula + code per step) for concept.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
concept
|
str
|
One of :func: |
required |
out
|
str | None
|
Output HTML path; defaults to |
None
|
open_browser
|
bool
|
Open the generated file in the default browser. |
True
|
Returns:
| Type | Description |
|---|---|
str
|
The path the HTML was written to. |
explore_concepts(out=None, open_browser=True)
¶
Build a searchable landing page linking to every :func:explain concept.
Each card links to explain_{key}.html alongside the generated page; run
:func:explain for a concept (or optimumai explain <concept>) to
inspect one directly.
list_explain_concepts()
¶
Return every concept key :func:explain accepts, sorted alphabetically.
editable_plot(expression='a*x^2 + b*x + c', params=None, xrange=(-5.0, 5.0), out=None)
¶
Build the interactive editor HTML.
If out is given, write the file and return its path; otherwise return the
HTML string. Parameters (single-letter names other than x) get sliders
unless you pass an explicit params mapping.
nn_playground(out=None)
¶
A TensorFlow-Playground-style neural-net playground, powered by OptiX.
Pick a 2-D dataset (XOR / circle / spiral), set the learning rate and hidden width, and train a tiny MLP while its decision boundary forms live. The math — seeded MLP init, forward pass, and backprop — is OptiX, a typed, unit-tested TypeScript kit compiled into OptimumAI. Self-contained, offline.
playground(name, out=None)
¶
Build the named playground ("attention", "kmeans", "astar", or "nn").
describe(data)
¶
Numpy summary statistics for a batch of numbers.
Returns n, mean, std, min, q1, median, q3,
max. Raises :class:ValueError on empty data.
plot_code(data, kind='bar', **opts)
¶
Return the exact, copy-paste-runnable matplotlib+numpy source.
The string always starts with import numpy as np and
import matplotlib.pyplot as plt, builds the same array Plot Studio
used, calls the matching plotting function, labels the axes/title, and
ends with plt.show().
kind is one of "bar", "hist", "scatter", "box",
"line", "pie", "violin". For "scatter", data may be a
list of (x, y) pairs or two equal-length sequences [xs, ys].
opts accepts title (all kinds) and bins ("hist" only,
default 10).
Raises :class:ValueError on empty data or an unknown kind.
plot_data(data, kind='bar', *, out='plot.png', **opts)
¶
Render data as kind with matplotlib and save it to out.
Returns the path to the saved PNG. See :func:plot_code for the
supported kind values, the accepted data shapes, and opts.
plot_studio_playground(out=None)
¶
Build the Plot Studio HTML playground: type numbers, get chart + code.
A textarea accepts numbers separated by commas, whitespace, or newlines
(for scatter, one x,y pair per line). A dropdown picks one of the
seven supported chart kinds. Three panels update live as you type: the
rendered chart (drawn on a <canvas> by hand-rolled JavaScript), a
numpy-style summary-statistics table, and the exact matplotlib+numpy
source that reproduces the current chart, with a Copy button.
Fully self-contained: inline CSS and vanilla JavaScript, no CDN, no server, works offline. Seeded with example data so it renders on open.
Submodule reference¶
optimumai.algebra
¶
Linear algebra with step-by-step explanations.
Matrix
¶
Vector
¶
A 1-D array with step-by-step explanations of its operations.
dot_trace(other)
¶
Build the full trace of self · other.
dot(other, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Inner product. Set explain=True to print the full trace.
norm_trace()
¶
Euclidean (L2) norm, ||a|| = √Σ aᵢ².
cosine_similarity_trace(other)
¶
Cosine similarity (a · b) / (||a|| ||b||) — the RAG workhorse.
optimumai.probability
¶
Probability building blocks used across modern AI.
softmax_trace(x, temperature=1.0)
¶
Build the full trace of softmax(x) with an optional temperature.
optimumai.calculus
¶
Calculus: numeric derivatives, gradients, and the chain rule.
chain_rule_trace(x=1.5)
¶
Demonstrate the chain rule on f(x) = tanh(x²) — exact vs numeric.
dy/dx = tanh'(x²) · d(x²)/dx = (1 − tanh²(x²)) · 2x.
derivative_trace(f, x, h=1e-05, label='f')
¶
Central-difference derivative of a scalar function at x.
gradient(f, point, h=1e-05, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Numeric gradient of f at point.
gradient_trace(f, point, h=1e-05)
¶
Numeric gradient (vector of partial derivatives) of f at point.
optimumai.autograd
¶
A tiny scalar autograd engine (micrograd-style) with a traceable backward pass.
Value
¶
A differentiable scalar node in an autograd graph.
zero_grad()
¶
Reset the gradient of every node reachable from this one.
backward()
¶
Run reverse-mode autodiff: seed grad=1 and apply the chain rule.
backward_trace()
¶
Like :meth:backward, but record how each gradient is produced.
For every node processed (output → inputs) we snapshot its children's
gradients before and after applying the local _backward, so the delta
we log is the exact contribution the chain rule pushed to each child.
backprop(explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Convenience: run backward, optionally rendering the chain-rule trace.
optimumai.optimization
¶
Optimizers and the training loop that turns gradients into learning.
SGD
¶
Vanilla stochastic gradient descent: w ← w − lr · ∂L/∂w.
Adam
¶
Adam: per-parameter adaptive steps from 1st/2nd gradient moments.
m ← β₁m + (1−β₁)g, v ← β₂v + (1−β₂)g² m̂ = m/(1−β₁ᵗ), v̂ = v/(1−β₂ᵗ) w ← w − lr · m̂ / (√v̂ + ε)
descent_demo(optimizer='adam', steps=50)
¶
Minimize the bowl f(x, y) = (x−3)² + (y+1)² from the origin.
The global minimum is at (3, −1) with loss 0 — an easy target to watch
the optimizer walk toward.
minimize(loss_fn, params, optimizer, steps=50, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Minimize loss_fn over params and return the final parameter values.
minimize_trace(loss_fn, params, optimizer, steps=50)
¶
Run steps optimization iterations, recording the loss each time.
loss_fn must rebuild a scalar-:class:Value loss from the (persistent)
params on every call. Each iteration recomputes the graph, backprops, and
lets the optimizer nudge the parameters.
optimumai.neural_networks
¶
Neural networks built on the scalar autograd engine.
MLP
¶
A stack of dense layers. Hidden layers use activation; output is linear.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n_in
|
int
|
Number of input features. |
required |
n_outs
|
Sequence[int]
|
Sizes of each layer, e.g. |
required |
activation
|
str
|
Hidden-layer nonlinearity ( |
'tanh'
|
seed
|
int
|
Seed for reproducible weight initialization. |
0
|
Layer
¶
A dense layer: n_out neurons, each seeing all inputs.
Neuron
¶
A single neuron: activation(Σ wᵢxᵢ + b).
forward_backward_trace(mlp, x, target)
¶
One example: forward to a prediction, squared-error loss, then backprop.
train(steps=100, lr=0.05, seed=0, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Run the training demo; returns the final parameter vector.
train_demo(steps=100, lr=0.05, seed=0)
¶
Train a 3→4→4→1 MLP on the toy set and watch the loss fall.
optimumai.transformers
¶
Transformer math, explained one stage at a time.
Attention
¶
Bases: BaseOp
Single-head scaled dot-product attention.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
d_k
|
int | None
|
Key/query dimension used for the |
None
|
TransformerBlock
¶
Bases: BaseOp
One pre-norm transformer decoder block (attention + feed-forward).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
d_model
|
int
|
Model / residual-stream dimension. |
required |
n_heads
|
int
|
Number of attention heads (must divide |
required |
d_ff
|
int | None
|
Hidden width of the feed-forward network (default |
None
|
MultiHeadAttention
¶
Bases: BaseOp
Multi-head scaled dot-product self-attention.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n_heads
|
int
|
Number of parallel attention heads. |
required |
d_model
|
int
|
Model dimension; must be divisible by |
required |
positional_encoding(seq_len, d_model, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Return the sinusoidal PE matrix. Set explain=True to print the trace.
positional_encoding_trace(seq_len, d_model)
¶
Build the full trace of the sinusoidal positional encoding matrix.
optimumai.world_models
¶
World models — learning to predict in representation space (LeCun's JEPA).
JEPA
¶
Bases: BaseOp
A minimal Joint-Embedding Predictive Architecture.
Uses a fixed, seeded linear encoder f (shared by context and target) and
a linear predictor g living entirely in embedding space. Everything is
deterministic given seed so traces are reproducible.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_dim
|
int
|
Dimension of a raw input view (e.g. flattened patch). |
required |
embed_dim
|
int
|
Dimension of the abstract representation space. |
required |
seed
|
int
|
RNG seed for the encoder/predictor weights. |
0
|
forward(context, target, explain=False, level='engineer')
¶
Compute the energy. Set explain=True to print the four-stage trace.
demo(seed=0)
classmethod
¶
A tiny, reproducible example: two noisy views of one latent vector.
target = context + small noise mimics two augmentations of the same
image, so a well-behaved encoder keeps the energy small.
optimumai.interpretability
¶
Interpretability — why neurons are polysemantic (Anthropic's superposition).
superposition_trace(n_features, n_neurons, sparsity=0.7, seed=0)
¶
Build the full trace of a toy superposition encode/decode cycle.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n_features
|
int
|
Number of distinct features (must exceed |
required |
n_neurons
|
int
|
Number of neurons / activation dimensions. |
required |
sparsity
|
float
|
Fraction of features that are inactive on a given input. |
0.7
|
seed
|
int
|
RNG seed for the weight matrix and the sparse feature vector. |
0
|
optimumai.frontier
¶
Frontier concepts — how today's large models are actually built and run.
FlashAttention (IO-aware tiling + online softmax), quantization (int8/int4), LoRA (parameter-efficient fine-tuning), and DPO (preference alignment).
flash_attention_trace(Q, K, V, block_size=2)
¶
Build the full trace of tiled attention with online softmax.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
Q
|
Query matrix |
required | |
K
|
Key matrix |
required | |
V
|
Value matrix |
required | |
block_size
|
int
|
How many rows of K/V (and of Q) live in one SRAM tile. |
2
|
lora_trace(d_in=8, d_out=8, rank=2, seed=0)
¶
Build the full trace of a LoRA adapter over a frozen weight W₀.
Constructs a seeded frozen W₀ plus the LoRA factors A (Gaussian) and
B (zeros), shows that ΔW = B·A = 0 at init, counts the trainable
parameters saved, and demonstrates the forward pass y = (W₀ + B·A)·x.
dequantize(q, scale, zero_point)
¶
Rebuild approximate floats: x̂ = (q − zero_point)·scale.
quantize(x, bits=8, scheme='symmetric', granularity='tensor')
¶
Quantize x to integer codes.
Returns:
| Type | Description |
|---|---|
|
|
|
|
|
quantize_trace(x, bits=8, scheme='symmetric', granularity='tensor')
¶
Build the full trace of quantizing (and dequantizing) x.
dpo(prompt='Explain gravity.', beta=0.1, seed=0, explain=False, level=ExplainLevel.ENGINEER)
¶
Return the DPO loss for one preference triple. explain=True prints the trace.
dpo_trace(prompt='Explain gravity.', beta=0.1, seed=0)
¶
Build the full trace of the DPO loss on one toy preference triple.
Uses small seeded per-token log-probabilities for a policy model and a frozen reference model over a chosen and a rejected response, then walks through the implicit rewards, the preference margin, and the final DPO loss.
optimumai.embeddings
¶
Embeddings — turning discrete tokens into dense vectors.
embedding_lookup(tokens, dim=4, seed=0, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Look up embedding rows for tokens. Set explain=True to print it.
embedding_lookup_trace(tokens, dim=4, seed=0)
¶
Build the full trace of looking up embedding rows for tokens.
nearest_neighbors(query, vocab, dim=4, seed=0, k=3, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Return the top-k cosine neighbours of query within vocab.
nearest_neighbors_trace(query, vocab, dim=4, seed=0, k=3)
¶
Rank vocab by cosine similarity to query in embedding space.
optimumai.rag
¶
Retrieval-augmented generation, traced end to end.
RAGPipeline
¶
Bases: BaseOp
A minimal, deterministic retrieval-augmented-generation pipeline.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
corpus
|
list[str] | None
|
The documents to retrieve from. Defaults to five ML sentences. |
None
|
dim
|
int
|
Embedding dimension for the bag-of-hashed-words vectors. |
16
|
seed
|
int
|
Base seed; each word gets its own vector from |
0
|
rag_flow(query=_DEFAULT_QUERY, k=2, out='rag_explainer.html')
¶
Build a RAG :class:~optimumai.core.flow_trace.FlowTrace and render it
to a self-contained D3 + KaTeX HTML file.
Parameters¶
query:
The question to retrieve context for. Cosine scores are computed from
the real :class:~optimumai.rag.pipeline.RAGPipeline — not toy values.
k:
Top-k chunks to retrieve.
out:
Output HTML path. Defaults to "rag_explainer.html" in the current
directory.
Returns¶
str The path to the generated HTML file.
Examples¶
CLI::
optimumai flow rag
optimumai flow rag --out my_rag.html
Python::
from optimumai.rag.flow import rag_flow
rag_flow(query="How tall is the Eiffel Tower?", out="rag.html")
build_rag_trace(query=_DEFAULT_QUERY, k=2, rerank_noise=0.04, seed=0)
¶
Build a :class:FlowTrace for a RAG pipeline run.
Parameters¶
query: The user question to retrieve context for. k: Top-k chunks to retrieve. rerank_noise: Small perturbation added to cosine scores to simulate a cross-encoder reranker (real re-ranking shifts scores without changing the embedding). seed: Random seed for the embedding model (for reproducibility).
Returns¶
FlowTrace
A validated trace with real cosine similarity scores computed by the
actual :class:~optimumai.rag.pipeline.RAGPipeline.
optimumai.diffusion
¶
Diffusion — forward noising and the reverse denoising idea.
forward_diffusion(x0, timesteps=10, beta_start=0.0001, beta_end=0.2, seed=0, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Noise x0 to timestep T. Set explain=True to print the trace.
forward_diffusion_trace(x0, timesteps=10, beta_start=0.0001, beta_end=0.2, seed=0)
¶
Build the full trace of DDPM forward diffusion applied to x0.
optimumai.foundations
¶
Foundations of the AI stack — the math, frameworks, and hardware underneath.
Tensors & integration, the PyTorch and JAX programming models (autograd, grad/jit/vmap), and the systems layer: the CUDA execution & memory model, tiled matmul kernels, the KV cache, and a VRAM budget calculator.
tiled_matmul(M=64, N=64, K=64, tile=16, explain=False, level=ExplainLevel.ENGINEER)
¶
Return C = A·B (seeded). Set explain=True to print the tiling trace.
tiled_matmul_trace(M=64, N=64, K=64, tile=16)
¶
Trace the naive → tiled → coalesced optimization of C = A · B.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
M, N, K
|
Dimensions of |
required | |
tile
|
int
|
Side length of the square tile loaded into shared memory. The global-memory reads drop by roughly this factor. |
16
|
The result is the real product matrix C (computed with a seeded NumPy
A @ B). meta carries naive_reads, tiled_reads and
reduction_factor (which equals tile).
memory_hierarchy_trace()
¶
Trace the GPU memory hierarchy from fastest/smallest to slowest/largest.
Walks registers → shared memory (L1) → global memory (VRAM) with approximate latencies and bandwidths, then explains arithmetic intensity (FLOP per byte) and why most kernels are memory-bound rather than compute-bound. The result is a small dict summarizing each level.
thread_hierarchy_trace(grid_dim=3, block_dim=4, thread_id=None)
¶
Trace the grid → block → warp → thread hierarchy of a GPU launch.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
grid_dim
|
int
|
Number of blocks in the (1-D) grid. |
3
|
block_dim
|
int
|
Number of threads per block. |
4
|
thread_id
|
int | None
|
A global thread index to decompose back into
|
None
|
grad_trace(f, x, h=1e-05)
¶
Show JAX's grad idea via a central-difference derivative of f at x.
We approximate f'(x) numerically here so the concept is visible without a tracer. JAX computes the exact derivative by tracing the pure function f and running reverse-mode autodiff over that trace — same answer, no round-off.
pytree_trace(tree)
¶
Flatten a nested dict/list/tuple tree to leaves + structure, then rebuild.
Every JAX transformation operates over pytrees: model parameters are nested
dicts, optimizer state is a pytree, grad returns a pytree matching the
inputs. Flatten → transform the flat leaves → unflatten is the pattern.
vmap_trace(f, batch)
¶
Demonstrate vmap: apply pure f across a batch axis with no Python loop.
We compute the loop version and the stacked version and show they agree. JAX's vmap does the same thing by pushing a batch dimension through the traced function, letting XLA vectorize it — no interpreter loop at runtime.
kv_cache_size(n_layers, n_heads, head_dim, seq_len, batch=1, bytes_per_elem=2, kv_heads=None)
¶
Return the KV cache size in bytes (no trace).
kv_heads defaults to n_heads (Multi-Head Attention); pass a smaller
number for Grouped-Query Attention or 1 for Multi-Query Attention.
kv_cache_trace(n_layers, n_heads, head_dim, seq_len, batch=1, bytes_per_elem=2, kv_heads=None)
¶
Build the full trace of a transformer KV cache's memory footprint.
integrate(f, a, b, method='trapezoid', n=100, explain=False, level=ExplainLevel.INTERMEDIATE)
¶
Return ∫ f from a to b. Set explain=True to print the trace.
integrate_trace(f, a, b, method='trapezoid', n=100)
¶
Estimate ∫ f(x) dx from a to b and trace the steps.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
f
|
Callable[[ndarray], ndarray]
|
A vectorized function accepting a NumPy array and returning one. |
required |
a
|
float
|
Lower bound of integration. |
required |
b
|
float
|
Upper bound of integration. |
required |
method
|
str
|
|
'trapezoid'
|
n
|
int
|
Number of grid intervals (trapezoid) or samples (Monte Carlo). |
100
|
tensor_intro_trace()
¶
Build a trace introducing tensors as n-dimensional arrays.
pytorch_autograd(explain=False, level=ExplainLevel.ENGINEER)
¶
Return the loss of the demo neuron. Set explain=True to print the trace.
pytorch_autograd_trace()
¶
Build a tiny neuron with :class:Value and narrate the PyTorch equivalents.
We compute y = tanh(w*x + b) and loss = (y - target)**2, then call
loss.backward(). Every step maps back to a line of PyTorch so you can see
that torch.autograd is just reverse-mode autodiff over a dynamic graph.
vram_estimate(params_billions, precision_bytes=2, training=True, optimizer='adam', batch=1, seq_len=2048, hidden=4096, n_layers=32, activation_factor=None)
¶
Return the estimated total VRAM in GB (no trace).
vram_trace(params_billions, precision_bytes=2, training=True, optimizer='adam', batch=1, seq_len=2048, hidden=4096, n_layers=32, activation_factor=None)
¶
Build the full trace of an LLM's VRAM budget, component by component.
optimumai.kernels
¶
GPU kernels from scratch — a pure-Python simulator, real kernels, and challenges.
KernelWorkbench
¶
GpuSim
¶
A serial simulator of a SIMT GPU launch.
Instantiate once per kernel run (or reuse and call :meth:reset). The core is
:meth:launch, which iterates every (block, thread) in the requested grid,
hands each a :class:ThreadCtx, and tallies memory traffic into self.stats.
reset()
¶
Clear all instrumentation so the simulator can be launched again.
launch(grid, block, kernel, *, flops=0)
¶
Run kernel once per thread in the grid × block launch.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
grid
|
int | tuple[int, int]
|
Blocks in the grid — |
required |
block
|
int | tuple[int, int]
|
Threads per block — |
required |
kernel
|
Callable[[ThreadCtx], None]
|
A function taking a single :class: |
required |
flops
|
int
|
Total floating-point ops the kernel performs, recorded for arithmetic-intensity reporting. Purely informational. |
0
|
Returns:
| Type | Description |
|---|---|
MemoryStats
|
|
GpuSim
|
Kernels write their output into caller-owned arrays, so the "output |
tuple[MemoryStats, GpuSim]
|
buffer" is whatever array the caller passed in and the kernel stored to. |
Threads are visited block-by-block; within a block they are grouped into
warps of :data:WARP_SIZE, and coalescing is judged per warp (consecutive
lanes → consecutive addresses). 2-D blocks are flattened in row-major
(ty, tx) order for warp grouping, matching CUDA's lane numbering.
MemoryStats
dataclass
¶
Instrumentation gathered over one :meth:GpuSim.launch.
Attributes:
| Name | Type | Description |
|---|---|---|
global_reads |
int
|
Number of scalar reads from global memory (VRAM). |
global_writes |
int
|
Number of scalar writes to global memory. |
shared_reads |
int
|
Number of scalar reads from shared memory (SRAM). |
shared_writes |
int
|
Number of scalar writes to shared memory. |
flops |
int
|
Floating-point operations declared for the launch (for intensity). |
coalesced |
bool | None
|
Whether every warp's global accesses hit consecutive addresses.
|
global_bytes
property
¶
Total bytes moved to/from global memory (fp32 accounting).
total_global
property
¶
Total scalar global-memory accesses (reads + writes).
arithmetic_intensity()
¶
FLOPs per global byte moved — the roofline x-axis.
Higher means more compute-bound (data is reused); lower means memory-bound
(the ALUs starve waiting on VRAM). Returns 0.0 when no bytes moved.
as_dict()
¶
A JSON-friendly summary suitable for a :class:Trace meta block.
ThreadCtx
¶
The per-thread handle a kernel uses to read/write memory and see its index.
A fresh context is created for every simulated thread by :meth:GpuSim.launch;
kernels never construct one directly. All memory helpers funnel through the
parent :class:GpuSim so a single :class:MemoryStats accumulates across the
whole grid.
gload(array, index)
¶
Read array[index] from global memory, counting the access.
gstore(array, index, value)
¶
Write value to array[index] in global memory, counting the access.
sload(buffer, index)
¶
Read buffer[index] from shared memory, counting the access.
sstore(buffer, index, value)
¶
Write value to buffer[index] in shared memory, counting the access.
available_backends()
cached
¶
Backends usable on this machine. Always includes "simulator".
backend_report()
¶
A human-readable summary of what's available and how to enable more.
list_kernels()
¶
Names of the available kernels, simple → complex.
run_kernel(name, explain=False, level='engineer')
¶
Run a kernel by name; explain=True renders the trace.
optimumai.ml
¶
Classical machine learning — the algorithms that predate (and still outlive) deep nets.
Every estimator here follows the same shape as the rest of OptimumAI: a
<name>_trace(...) function builds a full step-by-step
:class:~optimumai.core.trace.Trace with real numbers, and a thin class
(LinearRegression, KMeans, ...) wraps it in the familiar
fit/predict interface.
DecisionTree
¶
A shallow classification tree grown by greedy impurity-reduction splits.
Example
model = DecisionTree(max_depth=1).fit([[0], [1], [10], [11]], [0, 0, 1, 1]) model.predict([[0.5], [10.5]]) array([0, 1])
KMeans
¶
Lloyd's-algorithm k-means clustering.
Example
model = KMeans(k=2).fit([[0], [1], [10], [11]]) model.predict([[0.5], [10.5]]) array([0, 1])
KNN
¶
k-nearest-neighbors classifier: store the data, vote at prediction time.
Example
model = KNN(k=1).fit([[0], [1], [10], [11]], [0, 0, 1, 1]) model.predict([[0.5], [10.5]]) array([0, 1])
LinearRegression
¶
Ordinary least squares via the normal equation.
Example
model = LinearRegression().fit([[1], [2], [3]], [2, 4, 6]) model.predict([[4]]) array([8.])
LogisticRegression
¶
Binary logistic regression trained by batch gradient descent.
Example
model = LogisticRegression(lr=0.5, steps=200).fit([[0], [1], [2], [3]], [0, 0, 1, 1]) model.predict([[0.1], [2.9]]) array([0, 1])
fit(X, y)
¶
Run :attr:steps gradient-descent updates and store the final weights.
predict_proba(X)
¶
Return P(y=1|x) for new samples (must call :meth:fit first).
predict(X)
¶
Return the hard 0/1 label (threshold at 0.5).
explain(X, y, level=ExplainLevel.INTERMEDIATE)
¶
Fit, print the full trace (with :attr:steps gradient updates shown), return labels.
GaussianNB
¶
Gaussian Naive Bayes classifier.
Example
model = GaussianNB().fit([[0], [1], [10], [11]], [0, 0, 1, 1]) model.predict([[0.5], [10.5]]) array([0, 1])
PCA
¶
Principal Component Analysis via covariance eigendecomposition.
Example
model = PCA(n_components=1).fit([[0, 0], [1, 1], [2, 2], [3, 3]]) model.transform([[4, 4]]).shape (1, 1)
fit(X)
¶
Compute the mean and top principal components from X.
transform(X)
¶
Project new data onto the fitted principal components.
fit_transform(X)
¶
Fit on X and return the projected training data in one call.
explain(X, level=ExplainLevel.INTERMEDIATE)
¶
Fit, print the full center/covariance/eigen trace, and return the projection.
decision_tree_trace(X, y, max_depth=2, criterion='gini')
¶
Build a trace of the best-split search at the root, then grow a shallow tree.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
Features, shape |
required | |
y
|
Integer class labels, shape |
required | |
max_depth
|
int
|
Maximum depth of the fitted tree. |
2
|
criterion
|
str
|
|
'gini'
|
entropy_impurity(y)
¶
Shannon entropy −Σ pᵢ log₂(pᵢ) of the labels in y.
gini_impurity(y)
¶
Gini impurity 1 − Σ pᵢ² of the labels in y.
kmeans_trace(X, k=2, max_iters=10, init=None)
¶
Build the full Lloyd's-algorithm trace, clustering X into k groups.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
Data points, shape |
required | |
k
|
int
|
Number of clusters. |
2
|
max_iters
|
int
|
Safety cap on assign/update rounds. |
10
|
init
|
ndarray | None
|
Optional starting centroids, shape |
None
|
knn_trace(X_train, y_train, x_query, k=3)
¶
Build the full distance/vote trace classifying one query point.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X_train
|
Training features, shape |
required | |
y_train
|
Training labels, shape |
required | |
x_query
|
A single query point, shape |
required | |
k
|
int
|
Number of neighbors to vote. |
3
|
linear_regression_trace(X, y)
¶
Build the full normal-equation trace fitting ŷ = Xθ to (X, y).
logistic_regression_trace(X, y, lr=0.5, steps=3, theta0=None)
¶
Build a trace of a few gradient-descent steps fitting logistic regression.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
Feature matrix, shape |
required | |
y
|
Binary labels in |
required | |
lr
|
float
|
Learning rate for each gradient step. |
0.5
|
steps
|
int
|
Number of gradient-descent steps to trace (kept small — a real fit would run until the loss plateaus). |
3
|
theta0
|
ndarray | None
|
Optional starting weights (zeros by default). |
None
|
accuracy(y_true, y_pred)
¶
Fraction of predictions that exactly match the true labels.
accuracy_trace(y_true, y_pred)
¶
Build the trace of accuracy = fraction of exactly-correct predictions.
confusion_matrix(y_true, y_pred, labels=None)
¶
Confusion matrix (rows = true label, columns = predicted label).
confusion_matrix_trace(y_true, y_pred, labels=None)
¶
Build the trace of the confusion matrix (rows = true, columns = predicted).
mse(y_true, y_pred)
¶
Mean squared error between y_true and y_pred.
mse_trace(y_true, y_pred)
¶
Build the trace of mean squared error.
precision_recall_f1(y_true, y_pred, positive_label=1)
¶
Return {"precision": ..., "recall": ..., "f1": ...} for positive_label.
precision_recall_f1_trace(y_true, y_pred, positive_label=1)
¶
Build the trace of binary precision, recall, and F1 for positive_label.
r2_score(y_true, y_pred)
¶
R² (coefficient of determination) between y_true and y_pred.
r2_score_trace(y_true, y_pred)
¶
Build the trace of the R² (coefficient of determination) score.
roc_auc(y_true, y_scores)
¶
ROC-AUC for binary labels y_true given continuous y_scores.
roc_auc_trace(y_true, y_scores)
¶
Build the trace of ROC-AUC via the Mann-Whitney rank-sum formula.
y_true must be binary (0/1); y_scores are the model's continuous
scores or probabilities for the positive class.
naive_bayes_trace(X_train, y_train, x_query)
¶
Build the full prior/likelihood/posterior trace classifying one query point.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X_train
|
Training features, shape |
required | |
y_train
|
Training labels, shape |
required | |
x_query
|
A single query point, shape |
required |
pca_trace(X, n_components=1)
¶
Build the full center/covariance/eigendecomposition/project trace.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
X
|
Data, shape |
required | |
n_components
|
int
|
How many top principal components to keep ( |
1
|
optimumai.search
¶
Classical AI search — how agents find paths and choose moves.
Two families, four algorithms:
- Path-finding (:mod:
optimumai.search.uninformed, :mod:optimumai.search.informed) — :func:bfs, :func:dfs, :func:uniform_cost_search, :func:greedy_best_first, and :func:astarall search a :class:~optimumai.search.problem.Graphor :class:~optimumai.search.problem.GridWorldfor a route from a start state to a goal state. - Adversarial search (:mod:
optimumai.search.adversarial) — :func:minimaxand :func:alpha_betachoose a move in a two-player game tree, with alpha-beta pruning finding the identical value while visiting fewer nodes.
Every function has a *_trace counterpart (e.g. :func:astar_trace) that
returns a full :class:~optimumai.core.trace.Trace of the search: frontier
contents, expansion order, and (for informed/adversarial search) the g/h/f
values or the alpha-beta window at every step.
GameNode
dataclass
¶
A node in a small, explicit game tree.
A leaf has value set and no children; an internal node has
children and no value. player records whose turn it is to
move at this node ("max" or "min"), used only for the trace
narration — the recursion itself alternates automatically.
Attributes:
| Name | Type | Description |
|---|---|---|
name |
str
|
Short label for tracing, e.g. |
value |
float | None
|
The static evaluation, only set on leaves. |
children |
list[GameNode]
|
Child nodes reachable by one move, only set on internal nodes. |
player |
Literal['max', 'min']
|
Whose turn it is to move at this node. |
is_leaf
property
¶
True if this node has no children (a terminal position).
Graph
dataclass
¶
An explicit directed, weighted graph used as a search problem.
The adjacency structure is a plain dict[state, dict[neighbor, cost]] —
no external graph library needed. Edges are directed: add both
a -> b and b -> a for an undirected edge (see :meth:add_edge).
Attributes:
| Name | Type | Description |
|---|---|---|
adjacency |
dict[State, dict[State, float]]
|
|
heuristics |
dict[State, float]
|
Optional straight-line-style estimates |
Example
g = Graph() g.add_edge("A", "B", 1) g.add_edge("B", "C", 2) g.neighbors("A") {'B': 1} g.cost("B", "C") 2
add_edge(a, b, cost=1.0, bidirectional=True)
¶
Add an edge a -> b with the given cost (and the reverse, by default).
neighbors(state)
¶
Return {neighbor: edge_cost} reachable in one step from state.
cost(a, b)
¶
Return the edge cost of the direct move a -> b.
heuristic(state, goal)
¶
Return an estimate of the remaining cost from state to goal.
Falls back to 0 for states with no registered heuristic, which is always admissible (never overestimates) but uninformative.
states()
¶
All states that appear as either an edge source or destination.
GridWorld
dataclass
¶
A 2-D grid path-finding problem: 4-connected moves, optional walls.
Coordinates are (row, col) with row increasing downward, matching how
the grid is usually printed. Movement is restricted to the 4 orthogonal
neighbors (no diagonals), each costing 1 by default, so BFS already finds
shortest hop-count paths here — the more interesting question is how
much less work A* does than BFS/UCS once a heuristic is added.
Attributes:
| Name | Type | Description |
|---|---|---|
width |
int
|
Number of columns. |
height |
int
|
Number of rows. |
walls |
set[Coord]
|
Set of blocked |
Example
gw = GridWorld(width=3, height=3, walls={(1, 1)}) sorted(gw.neighbors((0, 1))) [(0, 0), (0, 2), (1, 1)]
in_bounds(state)
¶
True if state lies within the grid's rectangle.
is_free(state)
¶
True if state is in bounds and not a wall.
neighbors(state)
¶
Return {neighbor: 1.0} for the in-bounds, non-wall 4-neighbors.
Note the returned dict includes cells even when they happen to be walls filtered out — walls are excluded entirely, never returned.
cost(a, b)
¶
Cost of moving from a to an adjacent free cell b (always 1).
heuristic(state, goal, kind='manhattan')
¶
Estimate the remaining cost from state to goal.
manhattan (|dr| + |dc|) is admissible here because every move
costs exactly 1 and can only change row or column by 1 — the true
remaining cost can never be less than the Manhattan distance.
euclidean is also admissible (straight-line distance is never
longer than a path constrained to a grid) but less tight, so it
prunes less and A* typically expands more nodes with it than with
Manhattan on a 4-connected grid.
render(path=None)
¶
Render the grid as text: # walls, * path, . free cells.
alpha_beta(node, maximizing=True, alpha=float('-inf'), beta=float('inf'))
¶
Return the minimax value of node's game tree, computed with pruning.
alpha_beta_trace(node, maximizing=True, alpha=float('-inf'), beta=float('inf'))
¶
Build the full trace of minimax with alpha-beta pruning.
Returns the identical root value :func:minimax_trace would (see the
module docstring for why), while typically visiting far fewer nodes —
the trace records every prune with the (alpha, beta) window active
at the time.
minimax(node, maximizing=True)
¶
Return the minimax value of the root of node's game tree.
minimax_trace(node, maximizing=True)
¶
Build the full trace of plain minimax evaluation over a game tree.
astar(problem, start, goal)
¶
Return the optimal path from start to goal via A* (admissible h assumed).
astar_trace(problem, start, goal)
¶
Build the full trace of A* search (order by f(n) = g(n) + h(n)).
Optimal whenever problem.heuristic is admissible (never overestimates
the true remaining cost) — see the module docstring for the proof sketch.
greedy_best_first(problem, start, goal)
¶
Return a path from start to goal via greedy best-first search.
greedy_best_first_trace(problem, start, goal)
¶
Build the full trace of greedy best-first search (order by h(n) alone).
bfs(problem, start, goal)
¶
Return the fewest-edges path from start to goal (BFS).
bfs_trace(problem, start, goal)
¶
Build the full trace of breadth-first search from start to goal.
dfs(problem, start, goal)
¶
Return a path from start to goal found via DFS (not necessarily optimal).
dfs_trace(problem, start, goal)
¶
Build the full trace of depth-first search from start to goal.
uniform_cost_search(problem, start, goal)
¶
Return the minimum-cost path from start to goal (UCS / Dijkstra).
uniform_cost_search_trace(problem, start, goal)
¶
Build the full trace of uniform-cost search (Dijkstra to a single goal).
optimumai.rl
¶
Reinforcement learning — how an agent learns to act from reward alone.
Four building blocks, each handed off to the next:
- :mod:
optimumai.rl.mdp— the formal object (states, actions, transitions, rewards, discount) and the Bellman equation that "solves" it exactly when the environment model is known (value iteration, policy iteration). - :mod:
optimumai.rl.q_learning— model-free temporal-difference learning (Q-learning, SARSA) when the model is not known, only lived experience. - :mod:
optimumai.rl.policy_gradient— REINFORCE, which parameterizes the policy directly and differentiates through sampled actions via the score-function trick. - :mod:
optimumai.rl.ppo— PPO's clipped surrogate objective, the stabilized policy-gradient update that powers the reinforcement-learning stage of RLHF (contrasted with the closed-form DPO loss in :mod:optimumai.frontier.rlhf).
MDP
dataclass
¶
A finite, discounted Markov Decision Process.
Attributes:
| Name | Type | Description |
|---|---|---|
states |
list[str]
|
State labels, e.g. |
actions |
list[str]
|
Action labels, e.g. |
transition |
ndarray
|
|
reward |
ndarray
|
|
gamma |
float
|
Discount factor, |
terminal |
frozenset[int]
|
Optional set of state indices with no outgoing value (their value is fixed at 0 and they are excluded from the max/backup). |
expected_backup(values, s, a)
¶
The Bellman backup Σₛ' P(s'|s,a)[R(s,a,s') + γV(s')] for one (s, a).
BanditEnv
dataclass
¶
A stateless k-armed bandit: pick an arm, get a fixed reward.
Deterministic rewards keep the lesson focused on how REINFORCE shifts probability mass toward the best action, without the extra variance a stochastic reward would add on top of the policy's own sampling noise.
PPOSample
dataclass
¶
One (state, action) observation replayed from a rollout batch.
Attributes:
| Name | Type | Description |
|---|---|---|
old_logprob |
float
|
logπ_θ_old(a|s) — log-prob under the policy that generated this action (frozen for the whole batch). |
new_logprob |
float
|
logπ_θ(a|s) — log-prob under the policy currently being optimized (changes across optimizer steps/epochs). |
advantage |
float
|
Aₜ, how much better this action was than the state's baseline (e.g. from a critic / GAE — not recomputed here). |
GridWorld
dataclass
¶
A tiny, fully deterministic grid: step off the edge and you don't move.
The agent starts at start and the episode ends on reaching goal
(reward +goal_reward) or a cell in traps (reward trap_reward,
also terminal). Every other step costs step_reward (encourages short
paths). Deterministic transitions isolate the TD-learning behavior being
taught here from the extra variance a stochastic MDP would add.
step(pos, action)
¶
Apply action at pos; return (next_pos, reward, done).
policy_iteration(mdp, theta=1e-06, max_iterations=1000, explain=False, level=ExplainLevel.ENGINEER)
¶
Return V^π* for mdp. explain=True prints the full trace.
policy_iteration_trace(mdp, theta=1e-06, max_iterations=1000)
¶
Build the full trace of policy iteration converging to π* on mdp.
Alternates policy evaluation — solve the linear system V^π = R + γPV^π
exactly for the current (possibly suboptimal) policy — with policy
improvement — make the policy greedy w.r.t. that V^π — until the
policy stops changing. Guaranteed to converge in a finite number of
iterations for a finite MDP.
value_iteration(mdp, theta=1e-06, max_iterations=1000, explain=False, level=ExplainLevel.ENGINEER)
¶
Return V* for mdp. explain=True prints the full trace.
value_iteration_trace(mdp, theta=1e-06, max_iterations=1000)
¶
Build the full trace of value iteration converging to V* on mdp.
Repeatedly applies the Bellman optimality backup to every state, tracking
the largest change Δ each sweep, and stops once Δ < theta. The
greedy policy w.r.t. the converged values is then extracted.
reinforce(env=None, episodes=200, alpha=0.1, gamma=0.99, seed=0, explain=False, level=ExplainLevel.ENGINEER)
¶
Return the learned action-probability vector. explain=True prints the trace.
reinforce_trace(env=None, episodes=200, alpha=0.1, gamma=0.99, seed=0)
¶
Build the full trace of REINFORCE learning a softmax policy on env.
Each "episode" here is a single-step trajectory (state, action, reward) —
enough to show the sample → return → score-function-gradient → update
loop without the bookkeeping of multi-step return-to-go, which
:func:optimumai.rl.ppo.ppo_trace picks up with real advantages instead.
ppo_clip(batch=None, epsilon=0.2, explain=False, level=ExplainLevel.ENGINEER)
¶
Return the PPO clipped loss for batch. explain=True prints the trace.
ppo_clip_trace(batch=None, epsilon=0.2)
¶
Build the full trace of the PPO clipped surrogate objective on batch.
Walks every sample through the ratio, the unclipped and clipped terms,
whether/why clipping engaged, and the final batch-averaged loss
(L = -mean(L^CLIP) since optimizers minimize).
q_learning_trace(world=None, episodes=500, alpha=0.5, gamma=0.9, epsilon=0.1, seed=0)
¶
Build the full trace of tabular Q-learning on world (default: a 3x3 grid).
sarsa(world=None, episodes=500, alpha=0.5, gamma=0.9, epsilon=0.1, seed=0, explain=False, level=ExplainLevel.ENGINEER)
¶
Return the learned Q-table. explain=True prints the full trace.
sarsa_trace(world=None, episodes=500, alpha=0.5, gamma=0.9, epsilon=0.1, seed=0)
¶
Build the full trace of tabular SARSA on world (default: a 3x3 grid).
optimumai.nlp
¶
Classical NLP — the statistical machinery beneath modern language models.
Before attention, before embeddings even, NLP was tokenization, counting, and
dynamic programming. This package covers the fundamentals that still power
production search/retrieval systems and explain why the neural approach in
:mod:optimumai.transformers and :mod:optimumai.embeddings works:
- :mod:
optimumai.nlp.bpe— byte-pair encoding, how text becomes tokens - :mod:
optimumai.nlp.tfidf— TF-IDF, weighting words by how distinguishing they are - :mod:
optimumai.nlp.ngram— n-gram language models, counting your way to "next word" - :mod:
optimumai.nlp.edit_distance— Levenshtein distance via dynamic programming - :mod:
optimumai.nlp.word2vec— skip-gram, the seed idea behind every learned embedding
BPETokenizer
¶
Learn BPE merges from a corpus, then encode new words with them.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
num_merges
|
int
|
How many merge rules to learn during :meth: |
10
|
NGramModel
¶
A counting-based n-gram language model with add-k smoothing.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
n
|
int
|
Order of the model (2 = bigram, 3 = trigram, ...). Must be >= 1. |
2
|
k
|
float
|
Add-k smoothing constant ( |
1.0
|
vocab_size
property
¶
Number of distinct tokens seen during :meth:fit (including BOS/EOS).
fit(corpus)
¶
Count n-grams (and their contexts) over a list of sentences.
prob(context, word)
¶
Add-k smoothed P(word | context) for a (n-1)-length context.
sentence_log_prob(sentence)
¶
Sum of log P(w_i | context) over every n-gram in a padded sentence.
perplexity(corpus)
¶
PPL = exp(-(1/N) * sum of log-probs) over every n-gram in corpus.
TfidfVectorizer
¶
Fit a vocabulary + idf table on a corpus, then transform documents to vectors.
Mirrors the fit/transform shape of familiar vectorizer APIs, but stays
numpy + stdlib only, matching the "tf * smoothed-idf" formula documented
in :func:tfidf_trace.
SkipGramModel
¶
A tiny skip-gram word2vec model trained with full-softmax SGD.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
vocab
|
list[str]
|
The (deduplicated) vocabulary; row |
required |
dim
|
int
|
Embedding dimensionality. |
4
|
seed
|
int
|
Seed for the initial embedding matrices. |
0
|
predict(center)
¶
Return P(context | center) over the whole vocabulary.
step(center, context, lr=0.1)
¶
One full-softmax SGD step on a single (center, context) pair.
Returns the cross-entropy loss before the update.
most_similar(word, k=3)
¶
Rank the vocabulary by cosine similarity to word in w_in space.
bpe_trace(corpus, num_merges, encode_word)
¶
Train BPE on corpus and trace each merge round, then encode encode_word.
edit_distance_trace(a, b)
¶
Build the full DP-table + backtrace trace for the Levenshtein distance.
edit_script(a, b)
¶
Return the backtraced edit script turning a into b.
Each element is (op, from_char, to_char) with op in
{"match", "substitute", "delete", "insert"}; "-" marks the absent
side of an insert/delete.
ngram_trace(train_corpus, test_sentence, n=2, k=1.0)
¶
Fit an n-gram model on train_corpus and trace its PPL on test_sentence.
perplexity(model, corpus)
¶
Convenience wrapper: model.perplexity(corpus).
See :func:optimumai.evaluation.perplexity.perplexity for the general
formula applied to a bare list of per-token probabilities; this variant is
specific to a fitted :class:NGramModel scoring a corpus of sentences.
tfidf_trace(docs)
¶
Build the full TF, IDF, and weighted-matrix trace for a document set.
skipgram_pairs(corpus, window=1)
¶
Build every (center, context) pair within window of each center word.
word2vec_trace(corpus, window=1, dim=4, steps=20, lr=0.1, seed=0)
¶
Build (center, context) pairs, train a few SGD steps, and trace the whole thing.
optimumai.vision
¶
Computer vision fundamentals — how a CNN sees an image.
2D convolution (the local-pattern detector), pooling (translation-tolerant downsampling), Sobel edge detection (a fixed convolution finding boundaries), and a tiny CNN forward pass (conv -> ReLU -> pool -> flatten -> dense -> softmax) that ties the others together and headlines the tensor-shape story.
cnn_forward(image, kernel, dense_weights, dense_bias, pool_size=2)
¶
Run one image through conv -> ReLU -> max-pool -> flatten -> dense -> softmax.
cnn_forward_trace(image, kernel, dense_weights, dense_bias, pool_size=2)
¶
Build the full trace of a tiny CNN forward pass, headlined by shape flow.
dense(flat, weights, bias)
¶
A fully-connected layer: logits = weights @ flat + bias.
relu(z)
¶
Elementwise max(0, z) — zero out negative activations, keep the rest.
conv2d(x, w, stride=1, padding=0, mode='cross-correlate')
¶
Slide kernel w over image x and return the output feature map.
mode="cross-correlate" (default) is what every deep learning framework
calls "convolution": no kernel flip. mode="convolve" flips the kernel
180° first, matching the classical signal-processing definition.
conv2d_trace(x, w, stride=1, padding=0, mode='cross-correlate')
¶
Build the full trace of sliding w over x, window by window.
sobel_edges(x, padding=1)
¶
Return (magnitude, orientation) gradient maps for image x.
orientation is in radians, from atan2(Gy, Gx) (range (-π, π]).
sobel_edges_trace(x, padding=1)
¶
Build the full trace of Sobel edge detection: Gx, Gy, magnitude, orientation.
avg_pool2d(x, kernel_size=2, stride=None)
¶
Downsample x by averaging each kernel_size×kernel_size window.
max_pool2d(x, kernel_size=2, stride=None)
¶
Downsample x by taking the max of each kernel_size×kernel_size window.
pool2d_trace(x, kernel_size=2, stride=None, how='max')
¶
Build the full trace of pooling x, window by window.
optimumai.evaluation
¶
LLM evaluation — how do you score a model's output when there's no single right answer?
Four complementary angles: surface-overlap text metrics (BLEU/ROUGE/F1/EM),
intrinsic language-model quality (perplexity), whether a model's stated
confidence can be trusted (calibration/ECE), and whether a generated answer
stays grounded in its source (a hallucination/faithfulness heuristic — see
:mod:optimumai.evaluation.hallucination for why that last one is an
educational proxy, not a solved problem).
ece(confidences, correct, n_bins=5)
¶
Return the Expected Calibration Error in [0, 1] (lower is better calibrated).
ece_trace(confidences, correct, n_bins=5)
¶
Build the full ECE trace: bin the predictions, then show each bin's confidence-vs-accuracy gap.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
confidences
|
Sequence[float]
|
The model's stated confidence for each prediction, in |
required |
correct
|
Sequence[bool]
|
Whether each prediction was actually correct. |
required |
n_bins
|
int
|
Number of equal-width reliability bins (default 5, i.e. 20%-wide bins). |
5
|
faithfulness_score(answer, context)
¶
Return the claim-overlap faithfulness heuristic in [0, 1] (see module docstring).
faithfulness_trace(answer, context)
¶
Build the full trace: which answer tokens are supported by the context, and the score.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
answer
|
str
|
The model's generated answer to score. |
required |
context
|
str
|
The source text the answer was supposed to be grounded in (e.g. retrieved documents, a system prompt's reference material). |
required |
unsupported_spans(answer, context)
¶
Return the answer's content tokens absent from the context's vocabulary.
These are the heuristic's flagged candidates for hallucinated content — treat them as "worth checking," not as confirmed fabrications (see module docstring).
perplexity_trace(probs)
¶
Build the full trace: per-token surprisal → cross-entropy → perplexity.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
probs
|
Iterable[float]
|
The probability the model assigned to the actual next token
at each position (i.e. |
required |
bleu(candidate, reference, max_n=4)
¶
Return the BLEU score of candidate against reference in [0, 1].
bleu_trace(candidate, reference, max_n=4)
¶
Build the full BLEU trace: per-order clipped precision, then the brevity penalty.
exact_match(candidate, reference)
¶
1.0 if the normalized token sequences are identical, else 0.0 (SQuAD-style).
rouge_l(candidate, reference)
¶
Return the ROUGE-L F-measure in [0, 1].
rouge_l_trace(candidate, reference)
¶
Build the ROUGE-L trace: LCS length → recall/precision/F-measure.
rouge_n(candidate, reference, n=1)
¶
Return ROUGE-N recall (the conventionally reported value) in [0, 1].
rouge_n_trace(candidate, reference, n=1)
¶
Build the ROUGE-N trace: recall/precision/F1 of order-n n-gram overlap.
token_f1(candidate, reference)
¶
Return the token-level F1 of candidate against reference in [0, 1].
token_f1_trace(candidate, reference)
¶
Build the token-F1 trace: bag-of-tokens precision/recall/F1 (SQuAD-style).
optimumai.prompting
¶
Prompt-engineering patterns — constructed and explained offline.
Every pattern in this package builds a prompt deterministically, step by
step, with no live LLM calls: the point is to see exactly how the prompt
string is assembled and why the pattern helps (or where it fails), not to
call out to a model. Each submodule exposes a <name>_trace function
(returns a :class:~optimumai.core.trace.Trace), a thin <name> wrapper
(returns the assembled prompt string, or renders the trace if
explain=True), and a demo function for the curriculum/CLI.
chain_of_thought_trace(task, examples=None, trigger=DEFAULT_TRIGGER)
¶
Build a CoT prompt: optional worked exemplars, task, reasoning scaffold, answer slot.
few_shot_trace(task, examples, instruction='')
¶
Build a few-shot prompt: instruction + K exemplars + query, one exemplar at a time.
react_trace(task, tool_name='search', tool_description='search[query] — looks up query and returns a short snippet.')
¶
Build a ReAct prompt: tool spec, task, and one Thought/Action/Observation cycle.
self_consistency_trace(task, sampled_answers=None)
¶
Build the shared CoT prompt, then tally N deterministic toy sampled answers.
structured_output_trace(task, schema, example_completion=None)
¶
Build a schema-constrained prompt and validate a toy completion against it.
zero_shot_trace(task, instruction='Complete the task below.', role=_DEFAULT_ROLE)
¶
Build a zero-shot prompt: role + instruction + task, step by step.
optimumai.augmented_rnns
¶
Attention and Augmented Recurrent Neural Networks.
Based on distill.pub's Attention and Augmented Recurrent Neural Networks
(2016): three RNN-era ideas that gave networks capabilities beyond a fixed
hidden state — attention as differentiable memory access, Neural Turing
Machines' external read/write memory, and Adaptive Computation Time's learned,
variable compute. All three converge on the same core trick (score, softmax,
weighted blend) that later became transformer attention; see
:mod:optimumai.transformers.attention for that descendant.
NTMMemory
¶
A tiny Neural Turing Machine memory bank with content-based addressing.
Wraps :func:ntm_read / :func:ntm_write as stateful operations over a
fixed-size memory matrix, mirroring how a real NTM controller would issue
a read followed by a write each timestep.
adaptive_computation_time(halting_logits, eps=0.01)
¶
Run ACT's halting mechanism over a sequence of ponder-step logits.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
halting_logits
|
Iterable[float]
|
Raw (pre-sigmoid) halting score at each ponder step, in order. Must be non-empty. |
required |
eps
|
float
|
Small tolerance; pondering stops once the cumulative halting
probability reaches |
0.01
|
Returns:
| Type | Description |
|---|---|
dict[str, float | int | ndarray]
|
A dict with |
dict[str, float | int | ndarray]
|
including the halt step), |
dict[str, float | int | ndarray]
|
|
dict[str, float | int | ndarray]
|
|
dict[str, float | int | ndarray]
|
( |
adaptive_computation_time_trace(halting_logits, eps=0.01)
¶
Build the full trace of ACT's ponder-and-halt mechanism.
Shows the per-step halting probabilities, the cumulative sum, the step at which pondering halts, the leftover remainder, and the resulting ponder cost (the term added to the training loss).
attention_read(query, memory)
¶
Read a weighted blend of memory rows using content-based attention.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
query
|
ndarray
|
A 1-D vector of shape |
required |
memory
|
ndarray
|
A 2-D array of shape |
required |
Returns:
| Type | Description |
|---|---|
ndarray
|
The blended read vector, shape |
attention_read_trace(query, memory)
¶
Build the full trace of a content-based memory read.
Shows the relevance scores, the softmax attention weights (which sum to 1), and the resulting blended read vector.
ntm_read(key, memory, beta=1.0)
¶
Content-addressed read: blend memory rows by cosine similarity to key.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
key
|
ndarray
|
1-D vector of shape |
required |
memory
|
ndarray
|
2-D array of shape |
required |
beta
|
float
|
Sharpness ("key strength"); larger values make addressing more peaked around the best-matching slot. |
1.0
|
Returns:
| Type | Description |
|---|---|
ndarray
|
The read vector, shape |
ntm_trace(memory, read_key, write_key, erase, add, beta=1.0)
¶
Build the full trace of an NTM read followed by a write.
Shows the content-based addressing weights for both operations, the read vector, and the memory bank before/after the write.
ntm_write(memory, key, erase, add, beta=1.0)
¶
Content-addressed write: erase then add, blended by addressing weights.
Mᵢ ← Mᵢ · (1 − wᵢ · erase) + wᵢ · add for every slot i.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
memory
|
ndarray
|
2-D array of shape |
required |
key
|
ndarray
|
1-D vector of shape |
required |
erase
|
ndarray
|
1-D vector of shape |
required |
add
|
ndarray
|
1-D vector of shape |
required |
beta
|
float
|
Addressing sharpness, as in :func: |
1.0
|
Returns:
| Type | Description |
|---|---|
ndarray
|
The new memory bank, shape |
optimumai.quiz
¶
Quiz / active-recall engine — test yourself, don't just reread.
Question
dataclass
¶
A single multiple-choice question tied to one curriculum lesson.
answer is the index into choices of the correct option, and
explanation is a one-sentence justification shown after answering.
Quiz
¶
A gradeable quiz over the questions for a single lesson.
questions
property
¶
The questions for this lesson, in bank order.
grade(answers)
¶
Grade answers against the correct indices.
A length mismatch is tolerated: any question without a corresponding answer (and any answer past the last question) is scored as wrong.
check(index, choice)
¶
Return whether choice is the correct option for question index.
QuizResult
dataclass
¶
The outcome of grading one quiz attempt.
per_question records, in order, whether each question was answered
correctly, so a caller can highlight exactly what to revisit.
score
property
¶
Fraction correct in [0.0, 1.0] (0.0 for an empty quiz).
all_questions()
¶
A flat list of every question across all lessons (for spaced repetition).
available_quizzes()
¶
Sorted lesson ids that currently have at least one question in the bank.
optimumai.review
¶
optimumai.llm
¶
Real token generation via local Ollama, Hugging Face, Anthropic, or a toy model.
optimumai.visualization
¶
Visualization: terminal traces, plus matplotlib graphs (the [viz] extra).
The plotting functions import matplotlib lazily, so importing this package never requires matplotlib — you only need it when you actually draw.
landscape_demo(out=None)
¶
The headline demo: gradient descent tumbling down the Rosenbrock banana.
plot_loss_landscape(func='bowl', out=None, start=(-1.8, 1.8), lr=0.02, steps=40, kind='both')
¶
Draw a loss landscape (contour, surface, or both) with a descent path.
plot_activation(name='gelu', xlim=(-6.0, 6.0), out=None)
¶
Plot an activation and its (numeric) derivative on a single axes.
plot_attention(text=None, out=None)
¶
Heatmap of scaled dot-product self-attention weights with token labels.
plot_embeddings(words=None, dim=16, out=None)
¶
Scatter seeded word embeddings after a numpy-SVD PCA down to 2-D.
plot_heatmap(matrix, title='', labels=None, out=None)
¶
imshow a 2-D array with a colorbar; annotate cells when the matrix is small.
plot_softmax_temperature(logits=(2.0, 1.0, 0.1), temperatures=(0.5, 1.0, 2.0, 5.0), out=None)
¶
Grouped bars of softmax(logits / T) — sharp at low T, flat at high T.
plot_training_curve(losses=None, out=None)
¶
Plot training loss vs iteration on a log-y axis (runs the demo if needed).
render_trace(trace, level=ExplainLevel.INTERMEDIATE, console=None)
¶
Pretty-print trace at the requested detail level.
optimumai.circuit
¶
The circuit view — render a computation graph like an electrical circuit.
Build a :class:FlowGraph from a :class:~optimumai.autograd.value.Value or a
user expression, then render it as interactive HTML, Graphviz DOT, or a terminal
table — with each node's forward data and backward gradient on the wires.
FlowGraph
dataclass
¶
A flattened, renderable view of a :class:Value computation DAG.
Attributes:
| Name | Type | Description |
|---|---|---|
nodes |
list[FlowNode]
|
Every node — value nodes and operator pseudo-nodes. |
edges |
list[tuple[str, str]]
|
Directed wires as |
FlowNode
dataclass
¶
A single node in a :class:FlowGraph.
Attributes:
| Name | Type | Description |
|---|---|---|
id |
str
|
Stable identity string ( |
label |
str
|
Human label — a variable name ( |
data |
float
|
Forward value (the number that flows forward through the wire). |
grad |
float
|
Backward gradient (∂L/∂node — flows backward). |
op |
str
|
The op that produced this value ( |
is_op |
bool
|
|
build_from_expression(expression, variables=None, run_backward=True)
¶
Compile a safe arithmetic expression into a Value graph + FlowGraph.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
expression
|
str
|
An arithmetic string using variable names, numbers, and the
operators |
required |
variables
|
dict[str, float] | None
|
Optional map of name → value. Names not present default to
|
None
|
run_backward
|
bool
|
If |
True
|
Returns:
| Type | Description |
|---|---|
Value
|
|
FlowGraph
|
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If the expression is empty, cannot be parsed, or contains disallowed syntax (anything beyond names / numbers / arithmetic). |
build_from_value(root)
¶
Flatten a :class:Value DAG into a :class:FlowGraph.
Walks root._topo() (dependencies before dependents). Every Value
becomes a value node. Every Value that was produced by an op (has both
_op and children) additionally gets an operator pseudo-node wired in
between it and its children — mirroring micrograd's draw_dot::
child ─┐
├─▶ (op) ─▶ result
child ─┘
Leaves (no _op / no children) get no op node.
to_dot(fg)
¶
Emit a Graphviz DOT description of fg (left-to-right layout).
Value nodes are records showing label | data d.dd | grad g.gg; operator
pseudo-nodes are small ellipses carrying the op symbol.
to_html(fg, out)
¶
Write a self-contained interactive circuit page to out; return path.
Uses vis-network from a CDN (single <script> tag — no other local
assets). Value nodes show label, data (blue) and grad (orange)
on multiple lines; operator pseudo-nodes are small circles. Edges are the
directed "wires". Physics is enabled for a self-organising layout.
to_terminal(fg, console=None)
¶
Print fg as "the circuit" in the terminal using Rich.
Nodes are listed in topological order (as stored). Each value row shows its
label, the producing op, its data (blue) and grad (orange), plus a
one-line summary of the wires flowing into it.