Interactive & Explained¶
Five interactive layers built on top of the core library:
- RAG flow diagram — a D3 + KaTeX pipeline explainer that proves the
FlowTraceschema generalises across concepts - Prompt engineering — offline, deterministic traces for every standard prompting pattern
- Augmented RNNs — the distill.pub lineage from soft attention to NTMs to adaptive computation
- Interactive playgrounds — self-contained, offline HTML widgets
- Concept Explorer — 30 AI/ML concepts as steppable DAGs with a formula and runnable code panel for every step
RAG flow diagram — optimumai.rag.flow¶
A Transformer-Explainer-style progressive pipeline diagram for
Retrieval-Augmented Generation. Built on the concept-agnostic FlowTrace
schema (optimumai.core.flow_trace) — the renderer never knows what "RAG"
means; it only reads nodes, edges, and steps.
Pipeline stages visualised:
| Step | Stage | What you see |
|---|---|---|
| 1 | Chunking | Document splits into 3 chunks; edges activate |
| 2 | Embed chunks | Each chunk mapped to a vector; formula rendered via KaTeX |
| 3 | Index | Vectors added to the vector store |
| 4 | Embed query | Query vector computed; formula $\vec{q} = \text{Embed}(\text{query})$ |
| 5 | Retrieve | Cosine scores computed (real values from RAGPipeline); top-k selected |
| 6 | Rerank | Cross-encoder reranking stage |
| 7 | Assemble | Top chunks concatenated into context |
| 8 | Generate | LLM produces answer conditioned on context |
optimumai flow rag # Eiffel Tower demo query
optimumai flow rag --query "What year was it built?" # custom query, real scores
optimumai flow rag --out rag.html
from optimumai.rag.flow import rag_flow
from optimumai.rag.trace import build_rag_trace
from optimumai.rag.explainer import render_flow_trace_html, RAG_LAYOUT
from optimumai.core.flow_trace import FlowTrace
# One-liner: build trace + render HTML
path = rag_flow(out="rag_explainer.html")
# Or inspect the trace programmatically
trace = build_rag_trace(query="How tall is the Eiffel Tower?", k=2)
assert isinstance(trace, FlowTrace)
problems = trace.validate() # [] — referential integrity checks pass
# Real cosine scores from RAGPipeline._embed()
print(trace.steps[4].metrics)
# {'chunk_0_score': 0.6734, 'chunk_1_score': 0.3087, 'chunk_2_score': 0.6847}
# Render any FlowTrace — the renderer is concept-agnostic
html = render_flow_trace_html(trace, RAG_LAYOUT, out="rag.html")
Why the abstraction holds: The JavaScript in the generated HTML never
imports optimumai.rag. It reads TRACE.nodes, TRACE.edges,
TRACE.steps[i].formula (→ KaTeX), TRACE.steps[i].metrics (→ table).
Produce a FlowTrace from value_iteration or quantization and this same
renderer draws it, unmodified.
Internet required for the RAG diagram
The RAG diagram loads D3 v7 and KaTeX from CDN (the other flow
diagrams — transformer, attention, tfidf, word2vec — are fully offline
inline SVG). For offline use, download the two CDN files and replace the
<script>/<link> tags in the generated HTML.
Prompt engineering — optimumai.prompting¶
Each pattern builds the prompt step by step and explains why it works and how it fails. No API key needed — these are deterministic offline traces.
| Pattern | Command | What it teaches |
|---|---|---|
zero-shot |
optimumai prompt zero-shot |
Role + instruction + task, no examples |
few-shot |
optimumai prompt few-shot |
In-context learning from K exemplars |
chain-of-thought |
optimumai prompt chain-of-thought |
Elicit reasoning before the final answer |
react |
optimumai prompt react |
Thought / Action / Observation with a tool |
self-consistency |
optimumai prompt self-consistency |
Sample N chains, majority-vote the answer |
structured-output |
optimumai prompt structured-output |
Constrain to a validated JSON schema |
from optimumai.prompting import (
zero_shot, few_shot, chain_of_thought,
react, self_consistency, structured_output
)
chain_of_thought(
"If a train travels 60 miles in 2 hours, what is its speed?",
explain=True
)
self_consistency(
"What is 2 + 2?",
sampled_answers=["4", "4", "5"],
explain=True
) # -> "4"
optimumai prompt zero-shot
optimumai prompt few-shot
optimumai prompt chain-of-thought
optimumai prompt react
optimumai prompt self-consistency
optimumai prompt structured-output
optimumai.prompting
¶
Prompt-engineering patterns — constructed and explained offline.
Every pattern in this package builds a prompt deterministically, step by
step, with no live LLM calls: the point is to see exactly how the prompt
string is assembled and why the pattern helps (or where it fails), not to
call out to a model. Each submodule exposes a <name>_trace function
(returns a :class:~optimumai.core.trace.Trace), a thin <name> wrapper
(returns the assembled prompt string, or renders the trace if
explain=True), and a demo function for the curriculum/CLI.
chain_of_thought_trace(task, examples=None, trigger=DEFAULT_TRIGGER)
¶
Build a CoT prompt: optional worked exemplars, task, reasoning scaffold, answer slot.
few_shot_trace(task, examples, instruction='')
¶
Build a few-shot prompt: instruction + K exemplars + query, one exemplar at a time.
react_trace(task, tool_name='search', tool_description='search[query] — looks up query and returns a short snippet.')
¶
Build a ReAct prompt: tool spec, task, and one Thought/Action/Observation cycle.
self_consistency_trace(task, sampled_answers=None)
¶
Build the shared CoT prompt, then tally N deterministic toy sampled answers.
structured_output_trace(task, schema, example_completion=None)
¶
Build a schema-constrained prompt and validate a toy completion against it.
zero_shot_trace(task, instruction='Complete the task below.', role=_DEFAULT_ROLE)
¶
Build a zero-shot prompt: role + instruction + task, step by step.
Augmented RNNs — optimumai.augmented_rnns¶
The pre-transformer ideas that made attention mainstream, traced from distill.pub's 2016 post.
Attention as differentiable memory read¶
Content-based soft attention: score each memory slot by cosine similarity to a query, softmax the scores, blend the slots — the direct ancestor of transformer attention.
from optimumai.augmented_rnns import attention_read
import numpy as np
memory = np.array([[1., 0., -1.], [0., 1., 0.], [-1., 0., 1.], [.5, .5, .5]])
query = np.array([1., 0., -1.])
attention_read(query, memory, explain=True)
Neural Turing Machine — NTM¶
A full NTM read/write head: cosine-addressed soft attention for reading, and an erase/add mechanism for writing to memory. The first neural system with an explicit external memory that could be addressed by content.
from optimumai.augmented_rnns import ntm_read, ntm_write
import numpy as np
memory = np.random.default_rng(0).normal(size=(8, 4))
key = np.array([0.5, -1., 0., 0.3])
ntm_read(key, memory, beta=3.0, explain=True)
Adaptive Computation Time — ACT¶
The model itself decides when to stop computing: it emits a halting
probability at each step, and stops when the cumulative probability exceeds
1 − ε. It pays a "ponder cost" for extra steps. This is the idea behind
variable-depth computation in PonderNet and Universal Transformers.
from optimumai.augmented_rnns import adaptive_computation_time
import numpy as np
halting_probs = np.array([0.5, 1.2, 2.0, -0.3, 3.0])
adaptive_computation_time(halting_probs, eps=0.01, explain=True)
optimumai.augmented_rnns
¶
Attention and Augmented Recurrent Neural Networks.
Based on distill.pub's Attention and Augmented Recurrent Neural Networks
(2016): three RNN-era ideas that gave networks capabilities beyond a fixed
hidden state — attention as differentiable memory access, Neural Turing
Machines' external read/write memory, and Adaptive Computation Time's learned,
variable compute. All three converge on the same core trick (score, softmax,
weighted blend) that later became transformer attention; see
:mod:optimumai.transformers.attention for that descendant.
NTMMemory
¶
A tiny Neural Turing Machine memory bank with content-based addressing.
Wraps :func:ntm_read / :func:ntm_write as stateful operations over a
fixed-size memory matrix, mirroring how a real NTM controller would issue
a read followed by a write each timestep.
adaptive_computation_time(halting_logits, eps=0.01)
¶
Run ACT's halting mechanism over a sequence of ponder-step logits.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
halting_logits
|
Iterable[float]
|
Raw (pre-sigmoid) halting score at each ponder step, in order. Must be non-empty. |
required |
eps
|
float
|
Small tolerance; pondering stops once the cumulative halting
probability reaches |
0.01
|
Returns:
| Type | Description |
|---|---|
dict[str, float | int | ndarray]
|
A dict with |
dict[str, float | int | ndarray]
|
including the halt step), |
dict[str, float | int | ndarray]
|
|
dict[str, float | int | ndarray]
|
|
dict[str, float | int | ndarray]
|
( |
adaptive_computation_time_trace(halting_logits, eps=0.01)
¶
Build the full trace of ACT's ponder-and-halt mechanism.
Shows the per-step halting probabilities, the cumulative sum, the step at which pondering halts, the leftover remainder, and the resulting ponder cost (the term added to the training loss).
attention_read(query, memory)
¶
Read a weighted blend of memory rows using content-based attention.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
query
|
ndarray
|
A 1-D vector of shape |
required |
memory
|
ndarray
|
A 2-D array of shape |
required |
Returns:
| Type | Description |
|---|---|
ndarray
|
The blended read vector, shape |
attention_read_trace(query, memory)
¶
Build the full trace of a content-based memory read.
Shows the relevance scores, the softmax attention weights (which sum to 1), and the resulting blended read vector.
ntm_read(key, memory, beta=1.0)
¶
Content-addressed read: blend memory rows by cosine similarity to key.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
key
|
ndarray
|
1-D vector of shape |
required |
memory
|
ndarray
|
2-D array of shape |
required |
beta
|
float
|
Sharpness ("key strength"); larger values make addressing more peaked around the best-matching slot. |
1.0
|
Returns:
| Type | Description |
|---|---|
ndarray
|
The read vector, shape |
ntm_trace(memory, read_key, write_key, erase, add, beta=1.0)
¶
Build the full trace of an NTM read followed by a write.
Shows the content-based addressing weights for both operations, the read vector, and the memory bank before/after the write.
ntm_write(memory, key, erase, add, beta=1.0)
¶
Content-addressed write: erase then add, blended by addressing weights.
Mᵢ ← Mᵢ · (1 − wᵢ · erase) + wᵢ · add for every slot i.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
memory
|
ndarray
|
2-D array of shape |
required |
key
|
ndarray
|
1-D vector of shape |
required |
erase
|
ndarray
|
1-D vector of shape |
required |
add
|
ndarray
|
1-D vector of shape |
required |
beta
|
float
|
Addressing sharpness, as in :func: |
1.0
|
Returns:
| Type | Description |
|---|---|
ndarray
|
The new memory bank, shape |
Interactive playgrounds — optimumai.visualization.playgrounds¶
Self-contained HTML files with inline vanilla JS — no server, no build, works
offline. Generate one with optimumai playground <name>:
Attention playground¶
Inspired by Transformer Explainer:
- Hover a query token → the attention heatmap updates live
- Drag the temperature slider → scores re-softmax instantly
Softmax playground¶
- Drag any logit slider left/right → the probability bars recompute instantly
- Shows how temperature changes the sharpness/flatness of the distribution
Backprop playground¶
Drag any input (a, b, c, f) in the expression (a·b + c)·f:
- Forward values update (blue)
- Gradients update (orange)
k-means playground¶
- Click anywhere on the canvas to add a data point
- Watch Lloyd's algorithm re-run: assign to nearest centroid → recompute centroids → repeat
- See which points swap cluster assignments with each step
A* playground¶
- Click to draw/erase walls on a grid
- A* expands the frontier toward the goal in real time
- Open/closed sets highlighted; path shown in green
optimumai.visualization.playgrounds
¶
Interactive playgrounds — poke a small model and watch it react, live.
Inspired by poloclub.github.io/transformer-explainer and
playground.tensorflow.org: each function here computes a small, deterministic
example in Python (numpy, seeded), embeds the numbers as JSON inside a single
.html file, and lets vanilla JavaScript own all of the interaction (canvas
drawing, sliders, click handlers). There is no server, no build step, and no
runtime Python dependency — a user opens the file in any browser, online or
off, and starts dragging things.
Three playgrounds ship today:
- :func:
transformer_attention_playground— hover a token to see who it attends to; drag a temperature slider to watch the attention distribution sharpen (low temperature) or flatten (high temperature). - :func:
kmeans_playground— click to drop 2-D points, then step or run Lloyd's algorithm and watch centroids chase their clusters. - :func:
astar_playground— draw walls on a grid, then watch A* search expand its frontier before tracing the shortest path.
Use :func:playground to build any of them by name (handy for a CLI).
transformer_attention_playground(text='the cat sat on the mat', out=None)
¶
Transformer-Explainer-style self-attention playground.
Splits text into whitespace tokens, builds small seeded random query
and key projections, and computes the raw self-attention score matrix
S = Q @ K.T / sqrt(d). Python computes all the numbers once (so the
page is fully deterministic); the browser then owns the interaction:
- hover or click a token to highlight the row of the score matrix that belongs to it — "each row is where one token looks";
- drag the temperature slider to recompute
softmax(S / temperature)live in JavaScript. Low temperature sharpens the distribution onto a single key; high temperature flattens it toward uniform attention.
kmeans_playground(out=None)
¶
Lloyd's k-means playground: click to add points, then step or run.
A blank canvas starts with a small seeded set of points (so the page is non-empty and deterministic on load). Buttons let you add random points, single-step Lloyd's algorithm, run it to convergence with an animation, or reset. All of the k-means math (assignment + centroid update + inertia) runs in JavaScript so it can animate; Python only supplies the seeded starting points.
astar_playground(out=None)
¶
A* pathfinding playground on a grid with a Manhattan heuristic.
Draws a grid canvas with a fixed start (green) and goal (red) cell.
Click or drag on the grid to toggle walls; the mode selector lets you
instead move the start or goal cell. Pressing "Run" executes A* (in
JavaScript, so the whole search — including the visited-node overlay and
the final path — can be drawn) with h(n) = |dx| + |dy|, then reports
the number of nodes expanded and the path length.
nn_playground(out=None)
¶
A TensorFlow-Playground-style neural-net playground, powered by OptiX.
Pick a 2-D dataset (XOR / circle / spiral), set the learning rate and hidden width, and train a tiny MLP while its decision boundary forms live. The math — seeded MLP init, forward pass, and backprop — is OptiX, a typed, unit-tested TypeScript kit compiled into OptimumAI. Self-contained, offline.
playground(name, out=None)
¶
Build the named playground ("attention", "kmeans", "astar", or "nn").
Concept Explorer — optimumai.visualization.explain¶
30 foundational AI/ML concepts, each rendered as an interactive DAG you
step through — every node lights up in order, and the side panel shows a
KaTeX-rendered formula and a runnable optimumai code snippet for
that exact step. Self-contained, offline HTML (D3 + dagre for layout, KaTeX
for math — no server, no build step).
optimumai explain # list all 30 concepts
optimumai explain attention # Q,K,V -> QKᵀ -> scale -> softmax -> weighted sum
optimumai explain backpropagation # forward pass, then the chain rule in reverse
optimumai explain adam_optimizer # 1st/2nd moments, bias correction, the update
optimumai explore # a searchable landing page linking all 30
from optimumai import explain, explore_concepts, list_explain_concepts
list_explain_concepts() # -> 30 concept keys, sorted
explain("kmeans_clustering") # -> writes explain_kmeans_clustering.html, opens it
explore_concepts() # -> writes explore.html, opens it
All 30 concepts, grouped by area:
| Area | Concepts |
|---|---|
| Core math | sum_and_dot_product, gradient, variance |
| Neural net building blocks | weights_bias_neuron, activation_functions, softmax, layer_normalization, dropout |
| Training | backpropagation, gradient_descent, adam_optimizer, adamw_optimizer, cross_entropy_loss |
| Classical ML | linear_regression, logistic_regression, bias_variance_tradeoff, pca, kmeans_clustering, supervised_ml, unsupervised_ml, model_drift |
| NLP & transformers | tokenizer, embedding_lookup, attention, transformer_block, kv_cache, tfidf |
| RL & agents | q_learning, reinforcement_learning_overview, multi_agentic_workflow |
Each concept in the CONCEPTS registry is a small computation graph (nodes +
edges) walked one step at a time; every step carries a narration, a KaTeX
formula, and a code snippet. explore_concepts() builds a landing page with a
live search box over all of them — click a card to launch that concept's
explainer.
optimumai.visualization.explain
¶
DAG-style concept explainers — formula + Python code, side by side.
Each concept in :data:CONCEPTS is a small computation graph (nodes + edges)
walked one step at a time. Every step carries a plain-language narration, a
KaTeX-rendered formula, and a real, runnable optimumai code snippet for
that exact operation — so the graph, the math, and the code stay in sync as
you step through.
>>> from optimumai.visualization.explain import explain
>>> explain("attention") # doctest: +SKIP
'explain_attention.html'
The page is a single self-contained HTML file (D3 + dagre + KaTeX from a
CDN; no server, no build step) — open it in any browser online. Use
:func:explore_concepts to build a searchable landing page linking to every
concept, and :func:list_explain_concepts to see what's available.
list_explain_concepts()
¶
Return every concept key :func:explain accepts, sorted alphabetically.
explain(concept, out=None, open_browser=True)
¶
Build an interactive DAG explainer (formula + code per step) for concept.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
concept
|
str
|
One of :func: |
required |
out
|
str | None
|
Output HTML path; defaults to |
None
|
open_browser
|
bool
|
Open the generated file in the default browser. |
True
|
Returns:
| Type | Description |
|---|---|
str
|
The path the HTML was written to. |
explore_concepts(out=None, open_browser=True)
¶
Build a searchable landing page linking to every :func:explain concept.
Each card links to explain_{key}.html alongside the generated page; run
:func:explain for a concept (or optimumai explain <concept>) to
inspect one directly.
Concept gallery — optimumai.visualization.gallery¶
Per-concept matplotlib plots and animated GIFs for all 21+ registered concepts.
Also reachable via the optimumai visualize <concept> --fmt png|gif registry.
optimumai visualize # list every concept + its formats
optimumai visualize attention --fmt png --out attn.png
optimumai visualize gradient_descent --fmt gif --out gd.gif
from optimumai.visualization.concepts import render_concept, list_concepts
list_concepts()
render_concept("attention", fmt="gif", out="attn.gif")
Needs pip install "optimumai[viz]".
optimumai.visualization.gallery
¶
A gallery of per-concept plots and GIFs for the v1.1 modules.
Where :mod:optimumai.visualization.plots and
:mod:optimumai.visualization.animate cover the v1.0 foundations (activations,
attention, embeddings, gradient descent, ...), this module does the same job
for v1.1: classical ML (:mod:optimumai.ml), search (:mod:optimumai.search),
reinforcement learning (:mod:optimumai.rl), vision (:mod:optimumai.vision),
and evaluation (:mod:optimumai.evaluation).
Every function computes real data by calling the actual v1.1 module (no
placeholder numbers), then renders it with matplotlib. matplotlib is an
optional dependency (the optimumai[viz] extra) so it is imported lazily
inside each function — the base package still imports fine without it. Static
plots save a PNG; animations save a GIF via matplotlib's PillowWriter.
plot_kmeans(out='kmeans.png')
¶
Plot 2-D points colored by their fitted k-means cluster, plus centroids.
Fits :class:optimumai.ml.KMeans on two synthetic blobs and scatters every
point in its assigned cluster's color, with an x marker at each final
centroid.
animate_kmeans(out='kmeans.gif', fps=2)
¶
Animate Lloyd's algorithm: points recolor and centroids move each iteration.
Re-runs the assign/update steps of :func:optimumai.ml.kmeans.kmeans_trace
by hand (small k and few points) so each frame is a real intermediate
state of the algorithm, not an interpolation.
plot_decision_boundary(out='decision_boundary.png')
¶
Plot a classifier's decision regions over a 2-D toy dataset.
Trains :class:optimumai.ml.LogisticRegression on two separable blobs,
scores a mesh grid covering the data, and shades the predicted region with
contourf behind the scattered training points.
plot_astar_grid(out='astar_grid.png')
¶
Plot an A* search: walls, start/goal, explored cells, and the final path.
Runs :func:optimumai.search.informed.astar_trace on a
:class:optimumai.search.problem.GridWorld with a wall forcing a detour,
then shades every expanded cell and overlays the reconstructed path.
animate_astar(out='astar.gif', fps=3)
¶
Animate A*'s frontier expanding cell by cell, then reveal the final path.
Uses the same :func:optimumai.search.informed.astar_trace expansion
order as :func:plot_astar_grid, revealing one more expanded cell per
frame and adding the reconstructed path in the last few frames.
plot_value_function(out='value_function.png')
¶
Plot a gridworld state-value heatmap with greedy-policy arrows.
Runs :func:optimumai.rl.mdp.value_iteration_trace on a small
deterministic gridworld MDP, then draws V* as a heatmap and overlays
an arrow at every non-terminal cell pointing in its greedy action.
animate_value_iteration(out='value_iteration.gif', fps=2)
¶
Animate the value heatmap converging sweep by sweep.
Re-runs the same Bellman backup as
:func:optimumai.rl.mdp.value_iteration_trace by hand, capturing V
after every sweep so each frame is a real intermediate value function.
plot_conv_feature_map(out='conv_feature_map.png')
¶
Plot an input image, a kernel, and the resulting feature map side by side.
Uses :func:optimumai.vision.conv2d with a vertical-edge kernel on a
small synthetic image (the same "half dark, half light" pattern used in
the module's own demo).
plot_calibration(out='calibration.png')
¶
Plot a reliability diagram: per-bin confidence vs accuracy, with the diagonal.
Runs :func:optimumai.evaluation.calibration.ece_trace on a fixed,
moderately overconfident set of predictions and draws a bar per bin
(accuracy) next to the perfectly-calibrated diagonal.
plot_ppo_clip(out='ppo_clip.png', epsilon=0.2)
¶
Plot the PPO clipped surrogate objective vs. the probability ratio r.
Sweeps r over a range and evaluates min(r·A, clip(r, 1-ε, 1+ε)·A)
directly (the same formula used inside
:func:optimumai.rl.ppo.ppo_clip_trace) for both A=+1 and A=-1 —
the classic clip diagram.