Skip to content

How Lens works ​

Every Lens workflow revolves around one selection: a point or region on a rendered target, with an optional note. Lens links the target to its producing cells, builds bounded notebook context, and captures an annotated image. A code-mode agent uses that evidence to inspect, change, and verify the notebook, then shows its work back in the notebook. This page follows the selection through that lifecycle.

The collaboration loop ​

  1. Mark the resultThe person selects evidence and describes what should change.
  2. Ground the requestLens connects the mark to its producing cell and graph context.
  3. Revise and verifyThe agent changes notebook code and checks the affected result.
  4. Review or reopenLens brings the result into view and preserves the selection for another pass.

Try the loop ​

Press Select, mark one bar, and add a note such as "Make this blue." The panel reads the same Lens state available to a code-mode agent. Click the API cards in order to send activity and review feedback back to the notebook.

1 Press Select2 Mark a bar3 Try the API cards4 Watch the notebook respond

Each card calls the method printed on it. A real agent keeps the activity handle while code mode inspects, edits, runs, and verifies the producing cells.

Targets and selections ​

A target is the selectable unit. By default every rendered notebook output is a target, identified by its cell ID. A cell that shows its code but renders no output is a target through that code, so a person can select the cell itself and ask for a change "in this cell." Authors can also declare regions of their own HTML as targets, with or without notebook inputs. Custom targets shows how.

A point or region narrows attention inside the target. Lens stores the geometry normalized to the target, so the selection survives resizing and reconnects when the same output renders again in the same browser document.

Each selection has a stable S<n> label that is never reused during the Lens instance, an optional note, producing-cell references, and image status. Open selections live in Open. The current selection is the one most recently created or activated. It is the likely referent when a person says "this" or "here." One request to an agent can refer to several selections.

What the agent receives ​

Lens.context() returns a detached LensContext. It describes one selection-state revision and does not change when the notebook or Lens state changes afterwards. Mutating calls take that revision as expected_revision, so an agent cannot act on stale attention by accident.

EvidenceQuestion it answersWhere
Selection referenceWhat did the person select and ask for?references
Target descriptionWhat name and rendering reference did the author supply?references
DOM hintWhich rendered element sat under the point or region?references
Graph contextWhich cells and control values produced the target?text
Selection imageWhat did the target look like when the selection was captured?images
Cell-output imageWhat does the producing cell's output look like now?cell_image()

references is JSON-safe and builds immediately. text renders on first read and is cached. It lists producing cells and their nearest upstream dependencies in dependency order, keeps at most 64 cells inside a shared 24,000-character source budget, and reports what it omitted. Safely displayable native marimo control values appear inline. Passwords, file payloads, custom controls, and opaque state appear as [redacted] or [unavailable].

The graph decides computational relevance. The image preserves capture-time visual focus. The DOM hint distinguishes nearby labels, rows, and containers. The agent still interprets what a chart mark or application object means. LensContext defines every field and bound.

Selection images ​

Lens starts image capture after it stores a selection. The image is an annotated PNG of the whole target with the S<n> marker drawn on it. Large or scrolled targets add a detail view at readable scale, and small targets include nearby context. The selection stays usable while capture is pending or after it fails.

StatusMeaningBytes in images
pendingLens stored the selection and is preparing its image.No
availableThe stored image matches the current point or region.Yes
outdatedThe marker moved after the image was captured.Yes, prior position
failedLens could not produce an image for the current selection.No

Moving or resizing a marker starts replacement capture and keeps the previous image as outdated until the new one succeeds. Deleting or resolving a selection releases its bytes.

A cell-output image is different: a fresh, unannotated PNG of one current notebook output that an agent requests through cell_image() after changing and running a cell. It is a one-use transfer for verification. Connect an agent shows the polling loop.

Activity, reveal, and resolve ​

Lens gives an agent three ways to show work in the notebook. Each serves a different stage and has a different lifetime.

StageOperationVisible resultState change
Workstart_activity()Marks the selected target or cell the agent is working on.None
Reviewreveal()Brings the target or cell into view with a label and message.None
Resolutionresolve()Shows an Addressed receipt.Open selections become History.

Activity is transient. start_activity() returns an ActivityHandle, and stop_activity(handle) clears it only while that handle still owns the presentation, so a delayed stop cannot erase newer work. Activity also ends when its duration expires, a later activity or reveal replaces it, or the Lens view tears down.

Reveal brings a result into view after verification. A single-step reveal holds for duration_ms or until dismissed. A sequence of 1–16 steps is a Trail: an ordered explanation attached to notebook cells that the person pages through at their own pace. Trails work without any selection, which makes them the tool for "walk me through this notebook." Reveals are transient and end when a referenced cell or its upstream inputs change.

Resolve is the revision-checked state change. It validates the whole batch of selection IDs, releases their images, appends one History entry per selection, and returns the new revision. A summary of what changed and how it was verified appears beside the original request. When a reveal precedes resolution, Lens shows the receipt after the reveal hold ends, so the person sees the result first and the acknowledgement second.

History and reopen ​

A History entry keeps metadata for one resolved selection: label, target, geometry, note, timestamps, resolution revision, and summary. It carries no image bytes. History holds the newest 64 entries within 64,000 bytes and lives as long as the Lens instance.

Addressed records that the agent returned the request for review. It does not claim the answer is right. Press Reopen on a History entry to restore the selection as current in Open, with a fresh image captured from the current target. The reopened selection carries previousResolution so the agent can see the earlier outcome. Reopen requires the target to be available in the same browser document.

Browser and Python ​

Lens uses anywidget to connect a browser view to a Python model in the kernel. Each side owns what it can see.

BrowserPython
Finds rendered outputs and handles pointer and keyboard inputStores Open selections and History
Positions markers, the dock, and agent feedbackReads the live marimo graph and builds LensContext
Captures selection and cell-output imagesValidates revisions, activity, reveal, and resolution

Compact selection records and explicit commands cross the widget connection. PNG bytes travel separately and never enter trait state, JSON references, or the standalone text.

Continue with Selections for the person's controls, or Connect an agent for the agent's side of the loop.