Skip to content

Interactive & Explained

Five interactive layers built on top of the core library:

  1. RAG flow diagram — a D3 + KaTeX pipeline explainer that proves the FlowTrace schema generalises across concepts
  2. Prompt engineering — offline, deterministic traces for every standard prompting pattern
  3. Augmented RNNs — the distill.pub lineage from soft attention to NTMs to adaptive computation
  4. Interactive playgrounds — self-contained, offline HTML widgets
  5. Concept Explorer — 30 AI/ML concepts as steppable DAGs with a formula and runnable code panel for every step

RAG flow diagram — optimumai.rag.flow

A Transformer-Explainer-style progressive pipeline diagram for Retrieval-Augmented Generation. Built on the concept-agnostic FlowTrace schema (optimumai.core.flow_trace) — the renderer never knows what "RAG" means; it only reads nodes, edges, and steps.

Pipeline stages visualised:

Step Stage What you see
1 Chunking Document splits into 3 chunks; edges activate
2 Embed chunks Each chunk mapped to a vector; formula rendered via KaTeX
3 Index Vectors added to the vector store
4 Embed query Query vector computed; formula $\vec{q} = \text{Embed}(\text{query})$
5 Retrieve Cosine scores computed (real values from RAGPipeline); top-k selected
6 Rerank Cross-encoder reranking stage
7 Assemble Top chunks concatenated into context
8 Generate LLM produces answer conditioned on context
optimumai flow rag                                     # Eiffel Tower demo query
optimumai flow rag --query "What year was it built?"  # custom query, real scores
optimumai flow rag --out rag.html
from optimumai.rag.flow import rag_flow
from optimumai.rag.trace import build_rag_trace
from optimumai.rag.explainer import render_flow_trace_html, RAG_LAYOUT
from optimumai.core.flow_trace import FlowTrace

# One-liner: build trace + render HTML
path = rag_flow(out="rag_explainer.html")

# Or inspect the trace programmatically
trace = build_rag_trace(query="How tall is the Eiffel Tower?", k=2)
assert isinstance(trace, FlowTrace)
problems = trace.validate()   # [] — referential integrity checks pass

# Real cosine scores from RAGPipeline._embed()
print(trace.steps[4].metrics)
# {'chunk_0_score': 0.6734, 'chunk_1_score': 0.3087, 'chunk_2_score': 0.6847}

# Render any FlowTrace — the renderer is concept-agnostic
html = render_flow_trace_html(trace, RAG_LAYOUT, out="rag.html")

Why the abstraction holds: The JavaScript in the generated HTML never imports optimumai.rag. It reads TRACE.nodes, TRACE.edges, TRACE.steps[i].formula (→ KaTeX), TRACE.steps[i].metrics (→ table). Produce a FlowTrace from value_iteration or quantization and this same renderer draws it, unmodified.

Internet required for the RAG diagram

The RAG diagram loads D3 v7 and KaTeX from CDN (the other flow diagrams — transformer, attention, tfidf, word2vec — are fully offline inline SVG). For offline use, download the two CDN files and replace the <script>/<link> tags in the generated HTML.


Prompt engineering — optimumai.prompting

Each pattern builds the prompt step by step and explains why it works and how it fails. No API key needed — these are deterministic offline traces.

Pattern Command What it teaches
zero-shot optimumai prompt zero-shot Role + instruction + task, no examples
few-shot optimumai prompt few-shot In-context learning from K exemplars
chain-of-thought optimumai prompt chain-of-thought Elicit reasoning before the final answer
react optimumai prompt react Thought / Action / Observation with a tool
self-consistency optimumai prompt self-consistency Sample N chains, majority-vote the answer
structured-output optimumai prompt structured-output Constrain to a validated JSON schema
from optimumai.prompting import (
    zero_shot, few_shot, chain_of_thought,
    react, self_consistency, structured_output
)

chain_of_thought(
    "If a train travels 60 miles in 2 hours, what is its speed?",
    explain=True
)

self_consistency(
    "What is 2 + 2?",
    sampled_answers=["4", "4", "5"],
    explain=True
)   # -> "4"
optimumai prompt zero-shot
optimumai prompt few-shot
optimumai prompt chain-of-thought
optimumai prompt react
optimumai prompt self-consistency
optimumai prompt structured-output

optimumai.prompting

Prompt-engineering patterns — constructed and explained offline.

Every pattern in this package builds a prompt deterministically, step by step, with no live LLM calls: the point is to see exactly how the prompt string is assembled and why the pattern helps (or where it fails), not to call out to a model. Each submodule exposes a <name>_trace function (returns a :class:~optimumai.core.trace.Trace), a thin <name> wrapper (returns the assembled prompt string, or renders the trace if explain=True), and a demo function for the curriculum/CLI.

chain_of_thought_trace(task, examples=None, trigger=DEFAULT_TRIGGER)

Build a CoT prompt: optional worked exemplars, task, reasoning scaffold, answer slot.

few_shot_trace(task, examples, instruction='')

Build a few-shot prompt: instruction + K exemplars + query, one exemplar at a time.

react_trace(task, tool_name='search', tool_description='search[query] — looks up query and returns a short snippet.')

Build a ReAct prompt: tool spec, task, and one Thought/Action/Observation cycle.

self_consistency_trace(task, sampled_answers=None)

Build the shared CoT prompt, then tally N deterministic toy sampled answers.

structured_output_trace(task, schema, example_completion=None)

Build a schema-constrained prompt and validate a toy completion against it.

zero_shot_trace(task, instruction='Complete the task below.', role=_DEFAULT_ROLE)

Build a zero-shot prompt: role + instruction + task, step by step.


Augmented RNNs — optimumai.augmented_rnns

The pre-transformer ideas that made attention mainstream, traced from distill.pub's 2016 post.

Attention as differentiable memory read

Content-based soft attention: score each memory slot by cosine similarity to a query, softmax the scores, blend the slots — the direct ancestor of transformer attention.

from optimumai.augmented_rnns import attention_read
import numpy as np

memory = np.array([[1., 0., -1.], [0., 1., 0.], [-1., 0., 1.], [.5, .5, .5]])
query = np.array([1., 0., -1.])
attention_read(query, memory, explain=True)
optimumai augrnn attention

Neural Turing Machine — NTM

A full NTM read/write head: cosine-addressed soft attention for reading, and an erase/add mechanism for writing to memory. The first neural system with an explicit external memory that could be addressed by content.

from optimumai.augmented_rnns import ntm_read, ntm_write
import numpy as np

memory = np.random.default_rng(0).normal(size=(8, 4))
key = np.array([0.5, -1., 0., 0.3])
ntm_read(key, memory, beta=3.0, explain=True)
optimumai augrnn ntm

Adaptive Computation Time — ACT

The model itself decides when to stop computing: it emits a halting probability at each step, and stops when the cumulative probability exceeds 1 − ε. It pays a "ponder cost" for extra steps. This is the idea behind variable-depth computation in PonderNet and Universal Transformers.

from optimumai.augmented_rnns import adaptive_computation_time
import numpy as np

halting_probs = np.array([0.5, 1.2, 2.0, -0.3, 3.0])
adaptive_computation_time(halting_probs, eps=0.01, explain=True)
optimumai augrnn act

optimumai.augmented_rnns

Attention and Augmented Recurrent Neural Networks.

Based on distill.pub's Attention and Augmented Recurrent Neural Networks (2016): three RNN-era ideas that gave networks capabilities beyond a fixed hidden state — attention as differentiable memory access, Neural Turing Machines' external read/write memory, and Adaptive Computation Time's learned, variable compute. All three converge on the same core trick (score, softmax, weighted blend) that later became transformer attention; see :mod:optimumai.transformers.attention for that descendant.

NTMMemory

A tiny Neural Turing Machine memory bank with content-based addressing.

Wraps :func:ntm_read / :func:ntm_write as stateful operations over a fixed-size memory matrix, mirroring how a real NTM controller would issue a read followed by a write each timestep.

read(key)

Content-addressed read from the current memory state.

write(key, erase, add)

Content-addressed write; updates and returns the new memory state.

adaptive_computation_time(halting_logits, eps=0.01)

Run ACT's halting mechanism over a sequence of ponder-step logits.

Parameters:

Name Type Description Default
halting_logits Iterable[float]

Raw (pre-sigmoid) halting score at each ponder step, in order. Must be non-empty.

required
eps float

Small tolerance; pondering stops once the cumulative halting probability reaches 1 - eps (or the logits run out).

0.01

Returns:

Type Description
dict[str, float | int | ndarray]

A dict with halting_probs (sigmoid of each logit up to and

dict[str, float | int | ndarray]

including the halt step), cumulative (running sum at each step),

dict[str, float | int | ndarray]

halt_step (1-based index of the step that triggered halting),

dict[str, float | int | ndarray]

remainder (weight given to the final step), and ponder_cost

dict[str, float | int | ndarray]

(halt_step + remainder).

adaptive_computation_time_trace(halting_logits, eps=0.01)

Build the full trace of ACT's ponder-and-halt mechanism.

Shows the per-step halting probabilities, the cumulative sum, the step at which pondering halts, the leftover remainder, and the resulting ponder cost (the term added to the training loss).

attention_read(query, memory)

Read a weighted blend of memory rows using content-based attention.

Parameters:

Name Type Description Default
query ndarray

A 1-D vector of shape (d,) — "what am I looking for?".

required
memory ndarray

A 2-D array of shape (n, d) — n memory slots of dimension d — "what is stored?".

required

Returns:

Type Description
ndarray

The blended read vector, shape (d,).

attention_read_trace(query, memory)

Build the full trace of a content-based memory read.

Shows the relevance scores, the softmax attention weights (which sum to 1), and the resulting blended read vector.

ntm_read(key, memory, beta=1.0)

Content-addressed read: blend memory rows by cosine similarity to key.

Parameters:

Name Type Description Default
key ndarray

1-D vector of shape (d,) used to address memory.

required
memory ndarray

2-D array of shape (n, d) — the memory bank.

required
beta float

Sharpness ("key strength"); larger values make addressing more peaked around the best-matching slot.

1.0

Returns:

Type Description
ndarray

The read vector, shape (d,).

ntm_trace(memory, read_key, write_key, erase, add, beta=1.0)

Build the full trace of an NTM read followed by a write.

Shows the content-based addressing weights for both operations, the read vector, and the memory bank before/after the write.

ntm_write(memory, key, erase, add, beta=1.0)

Content-addressed write: erase then add, blended by addressing weights.

Mᵢ ← Mᵢ · (1 − wᵢ · erase) + wᵢ · add for every slot i.

Parameters:

Name Type Description Default
memory ndarray

2-D array of shape (n, d) — the memory bank before the write.

required
key ndarray

1-D vector of shape (d,) used to address memory.

required
erase ndarray

1-D vector of shape (d,) in [0, 1] — what to clear.

required
add ndarray

1-D vector of shape (d,) — what to write in.

required
beta float

Addressing sharpness, as in :func:ntm_read.

1.0

Returns:

Type Description
ndarray

The new memory bank, shape (n, d).


Interactive playgrounds — optimumai.visualization.playgrounds

Self-contained HTML files with inline vanilla JS — no server, no build, works offline. Generate one with optimumai playground <name>:

Attention playground

Inspired by Transformer Explainer:

  • Hover a query token → the attention heatmap updates live
  • Drag the temperature slider → scores re-softmax instantly
optimumai playground attention

Softmax playground

  • Drag any logit slider left/right → the probability bars recompute instantly
  • Shows how temperature changes the sharpness/flatness of the distribution
optimumai playground softmax

Backprop playground

Drag any input (a, b, c, f) in the expression (a·b + c)·f:

  • Forward values update (blue)
  • Gradients update (orange)
optimumai playground backprop

k-means playground

  • Click anywhere on the canvas to add a data point
  • Watch Lloyd's algorithm re-run: assign to nearest centroid → recompute centroids → repeat
  • See which points swap cluster assignments with each step
optimumai playground kmeans

A* playground

  • Click to draw/erase walls on a grid
  • A* expands the frontier toward the goal in real time
  • Open/closed sets highlighted; path shown in green
optimumai playground astar

optimumai.visualization.playgrounds

Interactive playgrounds — poke a small model and watch it react, live.

Inspired by poloclub.github.io/transformer-explainer and playground.tensorflow.org: each function here computes a small, deterministic example in Python (numpy, seeded), embeds the numbers as JSON inside a single .html file, and lets vanilla JavaScript own all of the interaction (canvas drawing, sliders, click handlers). There is no server, no build step, and no runtime Python dependency — a user opens the file in any browser, online or off, and starts dragging things.

Three playgrounds ship today:

  • :func:transformer_attention_playground — hover a token to see who it attends to; drag a temperature slider to watch the attention distribution sharpen (low temperature) or flatten (high temperature).
  • :func:kmeans_playground — click to drop 2-D points, then step or run Lloyd's algorithm and watch centroids chase their clusters.
  • :func:astar_playground — draw walls on a grid, then watch A* search expand its frontier before tracing the shortest path.

Use :func:playground to build any of them by name (handy for a CLI).

transformer_attention_playground(text='the cat sat on the mat', out=None)

Transformer-Explainer-style self-attention playground.

Splits text into whitespace tokens, builds small seeded random query and key projections, and computes the raw self-attention score matrix S = Q @ K.T / sqrt(d). Python computes all the numbers once (so the page is fully deterministic); the browser then owns the interaction:

  • hover or click a token to highlight the row of the score matrix that belongs to it — "each row is where one token looks";
  • drag the temperature slider to recompute softmax(S / temperature) live in JavaScript. Low temperature sharpens the distribution onto a single key; high temperature flattens it toward uniform attention.

kmeans_playground(out=None)

Lloyd's k-means playground: click to add points, then step or run.

A blank canvas starts with a small seeded set of points (so the page is non-empty and deterministic on load). Buttons let you add random points, single-step Lloyd's algorithm, run it to convergence with an animation, or reset. All of the k-means math (assignment + centroid update + inertia) runs in JavaScript so it can animate; Python only supplies the seeded starting points.

astar_playground(out=None)

A* pathfinding playground on a grid with a Manhattan heuristic.

Draws a grid canvas with a fixed start (green) and goal (red) cell. Click or drag on the grid to toggle walls; the mode selector lets you instead move the start or goal cell. Pressing "Run" executes A* (in JavaScript, so the whole search — including the visited-node overlay and the final path — can be drawn) with h(n) = |dx| + |dy|, then reports the number of nodes expanded and the path length.

nn_playground(out=None)

A TensorFlow-Playground-style neural-net playground, powered by OptiX.

Pick a 2-D dataset (XOR / circle / spiral), set the learning rate and hidden width, and train a tiny MLP while its decision boundary forms live. The math — seeded MLP init, forward pass, and backprop — is OptiX, a typed, unit-tested TypeScript kit compiled into OptimumAI. Self-contained, offline.

playground(name, out=None)

Build the named playground ("attention", "kmeans", "astar", or "nn").


Concept Explorer — optimumai.visualization.explain

30 foundational AI/ML concepts, each rendered as an interactive DAG you step through — every node lights up in order, and the side panel shows a KaTeX-rendered formula and a runnable optimumai code snippet for that exact step. Self-contained, offline HTML (D3 + dagre for layout, KaTeX for math — no server, no build step).

optimumai explain                  # list all 30 concepts
optimumai explain attention        # Q,K,V -> QKᵀ -> scale -> softmax -> weighted sum
optimumai explain backpropagation  # forward pass, then the chain rule in reverse
optimumai explain adam_optimizer   # 1st/2nd moments, bias correction, the update
optimumai explore                  # a searchable landing page linking all 30
from optimumai import explain, explore_concepts, list_explain_concepts

list_explain_concepts()          # -> 30 concept keys, sorted
explain("kmeans_clustering")     # -> writes explain_kmeans_clustering.html, opens it
explore_concepts()               # -> writes explore.html, opens it

All 30 concepts, grouped by area:

Area Concepts
Core math sum_and_dot_product, gradient, variance
Neural net building blocks weights_bias_neuron, activation_functions, softmax, layer_normalization, dropout
Training backpropagation, gradient_descent, adam_optimizer, adamw_optimizer, cross_entropy_loss
Classical ML linear_regression, logistic_regression, bias_variance_tradeoff, pca, kmeans_clustering, supervised_ml, unsupervised_ml, model_drift
NLP & transformers tokenizer, embedding_lookup, attention, transformer_block, kv_cache, tfidf
RL & agents q_learning, reinforcement_learning_overview, multi_agentic_workflow

Each concept in the CONCEPTS registry is a small computation graph (nodes + edges) walked one step at a time; every step carries a narration, a KaTeX formula, and a code snippet. explore_concepts() builds a landing page with a live search box over all of them — click a card to launch that concept's explainer.

optimumai.visualization.explain

DAG-style concept explainers — formula + Python code, side by side.

Each concept in :data:CONCEPTS is a small computation graph (nodes + edges) walked one step at a time. Every step carries a plain-language narration, a KaTeX-rendered formula, and a real, runnable optimumai code snippet for that exact operation — so the graph, the math, and the code stay in sync as you step through.

>>> from optimumai.visualization.explain import explain
>>> explain("attention")               # doctest: +SKIP
'explain_attention.html'

The page is a single self-contained HTML file (D3 + dagre + KaTeX from a CDN; no server, no build step) — open it in any browser online. Use :func:explore_concepts to build a searchable landing page linking to every concept, and :func:list_explain_concepts to see what's available.

list_explain_concepts()

Return every concept key :func:explain accepts, sorted alphabetically.

explain(concept, out=None, open_browser=True)

Build an interactive DAG explainer (formula + code per step) for concept.

Parameters:

Name Type Description Default
concept str

One of :func:list_explain_concepts (case/dash/space-insensitive).

required
out str | None

Output HTML path; defaults to explain_{concept}.html.

None
open_browser bool

Open the generated file in the default browser.

True

Returns:

Type Description
str

The path the HTML was written to.

explore_concepts(out=None, open_browser=True)

Build a searchable landing page linking to every :func:explain concept.

Each card links to explain_{key}.html alongside the generated page; run :func:explain for a concept (or optimumai explain <concept>) to inspect one directly.


Per-concept matplotlib plots and animated GIFs for all 21+ registered concepts. Also reachable via the optimumai visualize <concept> --fmt png|gif registry.

optimumai visualize                              # list every concept + its formats
optimumai visualize attention --fmt png --out attn.png
optimumai visualize gradient_descent --fmt gif --out gd.gif
from optimumai.visualization.concepts import render_concept, list_concepts

list_concepts()
render_concept("attention", fmt="gif", out="attn.gif")

Needs pip install "optimumai[viz]".

optimumai.visualization.gallery

A gallery of per-concept plots and GIFs for the v1.1 modules.

Where :mod:optimumai.visualization.plots and :mod:optimumai.visualization.animate cover the v1.0 foundations (activations, attention, embeddings, gradient descent, ...), this module does the same job for v1.1: classical ML (:mod:optimumai.ml), search (:mod:optimumai.search), reinforcement learning (:mod:optimumai.rl), vision (:mod:optimumai.vision), and evaluation (:mod:optimumai.evaluation).

Every function computes real data by calling the actual v1.1 module (no placeholder numbers), then renders it with matplotlib. matplotlib is an optional dependency (the optimumai[viz] extra) so it is imported lazily inside each function — the base package still imports fine without it. Static plots save a PNG; animations save a GIF via matplotlib's PillowWriter.

plot_kmeans(out='kmeans.png')

Plot 2-D points colored by their fitted k-means cluster, plus centroids.

Fits :class:optimumai.ml.KMeans on two synthetic blobs and scatters every point in its assigned cluster's color, with an x marker at each final centroid.

animate_kmeans(out='kmeans.gif', fps=2)

Animate Lloyd's algorithm: points recolor and centroids move each iteration.

Re-runs the assign/update steps of :func:optimumai.ml.kmeans.kmeans_trace by hand (small k and few points) so each frame is a real intermediate state of the algorithm, not an interpolation.

plot_decision_boundary(out='decision_boundary.png')

Plot a classifier's decision regions over a 2-D toy dataset.

Trains :class:optimumai.ml.LogisticRegression on two separable blobs, scores a mesh grid covering the data, and shades the predicted region with contourf behind the scattered training points.

plot_astar_grid(out='astar_grid.png')

Plot an A* search: walls, start/goal, explored cells, and the final path.

Runs :func:optimumai.search.informed.astar_trace on a :class:optimumai.search.problem.GridWorld with a wall forcing a detour, then shades every expanded cell and overlays the reconstructed path.

animate_astar(out='astar.gif', fps=3)

Animate A*'s frontier expanding cell by cell, then reveal the final path.

Uses the same :func:optimumai.search.informed.astar_trace expansion order as :func:plot_astar_grid, revealing one more expanded cell per frame and adding the reconstructed path in the last few frames.

plot_value_function(out='value_function.png')

Plot a gridworld state-value heatmap with greedy-policy arrows.

Runs :func:optimumai.rl.mdp.value_iteration_trace on a small deterministic gridworld MDP, then draws V* as a heatmap and overlays an arrow at every non-terminal cell pointing in its greedy action.

animate_value_iteration(out='value_iteration.gif', fps=2)

Animate the value heatmap converging sweep by sweep.

Re-runs the same Bellman backup as :func:optimumai.rl.mdp.value_iteration_trace by hand, capturing V after every sweep so each frame is a real intermediate value function.

plot_conv_feature_map(out='conv_feature_map.png')

Plot an input image, a kernel, and the resulting feature map side by side.

Uses :func:optimumai.vision.conv2d with a vertical-edge kernel on a small synthetic image (the same "half dark, half light" pattern used in the module's own demo).

plot_calibration(out='calibration.png')

Plot a reliability diagram: per-bin confidence vs accuracy, with the diagonal.

Runs :func:optimumai.evaluation.calibration.ece_trace on a fixed, moderately overconfident set of predictions and draws a bar per bin (accuracy) next to the perfectly-calibrated diagonal.

plot_ppo_clip(out='ppo_clip.png', epsilon=0.2)

Plot the PPO clipped surrogate objective vs. the probability ratio r.

Sweeps r over a range and evaluates min(r·A, clip(r, 1-ε, 1+ε)·A) directly (the same formula used inside :func:optimumai.rl.ppo.ppo_clip_trace) for both A=+1 and A=-1 — the classic clip diagram.