Turn many sources into one precise, traceable skill
Not another pile of saved prompts. A place to turn overlapping source ideas into concise guidance you can inspect, reshape, and deliberately accept—without losing the words, differences, or provenance underneath.
SkillBenchLab puts the shared meaning first. Research brings proposals; the human owns the result. Its central promise is not “more sources must be better.” It is “you can see exactly what you are accepting, and why.”
Status: the warm-editorial prototype has been exercised in a real browser. Production source exists with known gaps. The actual browser → existing OMP conversation → real research → durable proposals chain is not yet verified. This page is a product pitch and an evidence-backed design, not a production launch announcement.
Real interface capture, not a concept render. The initial pin is an illustrative fixture—not a record of your acceptance.
01 · The unit of progress is a meaning, not a document
Reading ten related documents can leave you with ten copies of almost the same instruction. A flat summary has the opposite problem: it can erase the disagreement that mattered. SkillBenchLab is designed to hold both levels at once.
- The main surface:
- One compact essence row per shared meaning.
- The clearest useful wording, rather than a document title or a wall of excerpts.
- Recurrence sorted by distinct logical documents.
- Under each row:
- Exact quotations and separately labeled paraphrases.
- Source-specific flavours, context, uncertainty, and unresolved differences.
- Links back to the evidence that produced the interpretation.
- The intended payoff:
- Compare interpretations instead of rereading everything.
- Keep the distinctions worth keeping; remove repetition without erasing provenance.
- Produce guidance for a stated use and environment—not an allegedly universal truth.
The immediate audience is people researching, authoring, and adapting agent skills. Demand, time savings, and better agent outcomes have not been measured.
02 · Every essence has an inspectable underside
An essence is the editable synthesis. An impulse is an attributed piece of source evidence. A document observation records what was retrieved, from where, and at which version. These are different objects because they answer different questions.
Conceptual model. Example wording inside this diagram is illustrative, not an upstream quotation.
- Evidence trail:
- Source URL and title; version or content hash; retrieval time; license where known.
- Exact quote, extractor paraphrase, source-specific flavour, and context note.
- The research request responsible for a proposal.
- Editorial trail:
- Current working wording and its evidence.
- Differences still unresolved.
- Explicit accepted snapshots and their reviewed evidence.
- Three identities that must not be confused:
- Observation: a particular retrieval or version.
- Logical document: the work being counted.
- Lineage: a shared origin such as one author or repository.
This makes provenance inspectable. It does not prove that an interpretation is correct, that a license permits every intended reuse, or that the resulting skill improves performance.
Real prototype disclosure. The compact row is an entry point, not an excuse to discard the source material.
03 · Count recurrence without manufacturing confidence
A mirror is not another supporting document. A refetch is not another vote. Two useful quotations from one document are still one document for recurrence.
- The counting rule:
- Count distinct logical document identities attached to the essence.
- Preserve their individual observations and excerpts for inspection.
- Expose shared lineage separately; do not turn the count into a confidence score.
- The verified replay example:
- A further excerpt gives one essence four excerpts but still three documents.
- A known mirror/refetch does not increase recurrence.
- A genuinely additional document can increase the document count.
- The boundary:
- Supported identity normalization is not universal automatic mirror detection.
- Unknown identity and conflicting lineage need disclosure and reconciliation.
- The production mirror-reconciliation path is not yet exposed through the product.
04 · Review the wording and the evidence together
Acceptance is an editorial decision, not a side effect of research finishing or an arbitrary demand to change a few words.
- In the verified prototype:
- Compare the current proposal with the pinned snapshot before accepting.
- Accept an evidence-only revision without artificial wording edits.
- Reject reacceptance of an identical snapshot.
- Preserve accepted wording, evidence, and context while the working state changes.
- Restore historical working content without rewriting the current accepted pin.
- In the intended product contract:
- Research may propose and revise; it never impersonates human acceptance.
- New evidence goes into reviewable working state, not silently into an accepted export.
- A stronger claim requires an explicit new decision.
- Production parity is unfinished:
- Its text-only acceptance guard still blocks evidence-only acceptance.
- Snapshot context/caveat coverage and restoration behavior do not yet match the prototype.
05 · Research in the conversation you already have
The intended research engine is the existing Oh My Pi conversation, using its actual research tools and context—not a second chatbot hidden behind another API-key form.
Target workflow. Bridge code exists; the complete live same-conversation chain remains unproven.
- Start with context:
- The question, scope, intended use, and environment.
- Visible source, round, and time ceilings. Limits are ceilings, not quotas to fill.
- Let research do its part:
- Gather sources, retain attribution, and extract useful impulses.
- Propose shared essences; show uncertainty and source differences.
- Report coverage and incomplete work honestly.
- Keep control legible:
- Pause and cancel cooperatively at checkpoints, rather than promising instant tool interruption.
- Preserve already-applied proposals when a run is canceled.
- Reconcile interrupted or uncertain work before resubmitting.
- Divide authority deliberately:
- The agent looks up facts and proceeds on reversible choices the user has delegated.
- The human decides consequential tradeoffs and explicitly accepts guidance.
- Later answers can redirect the work; missing optional answers do not have to stop it.
The standalone demo instead replays four recorded checkpoints. It performs no network retrieval or inference. Its pause/resume controls a local timer; restart deduplicates previously applied replay proposals. These interactions demonstrate the interface contract, not the live engine.
06 · Shape the meaning, then export the reviewed result
Editorial work is not limited to accepting or rejecting a generated paragraph. Sometimes a row contains two ideas. Sometimes two rows are the same idea. Sometimes the right answer is to remove one.
- Verified prototype controls:
- Edit wording directly, with keyboard save and cancellation.
- Split a selected evidence subset when meanings deserve separate rows.
- Fuse selected rows when their meanings belong together.
- Remove with confirmation; undo structural work while it remains safe.
- Refuse the one-level undo after an intervening saved mutation rather than silently overwriting later work.
- Fusion and removal retire the original rows from active export; their immutable snapshots remain in page memory. A fused row starts unaccepted.
- Export contract demonstrated by the prototype:
- Default to active accepted snapshots, not working drafts.
- Respect a scoped selection.
- Include unaccepted working material only through an explicitly non-authoritative draft appendix.
- Keep provenance, attribution, license notes, and limitations visible in portable
SKILL.md.
The actual downloaded prototype file matched its preview: 4,635 UTF-8 bytes in the recorded exercise. Its pinned export remained byte-identical before and after incoming replay evidence. Production export currently handles drafts and context differently; this demonstration is not a claim of parity.
A concrete case: five documents, one shared lineage
The illustrative corpus is five engineering/productivity skill documents from Matt Pocock’s skills repository, pinned to commit 3cca18b368ae95cdbdebbff572ccafa662551015: Wayfinder, Grilling, Domain Modeling, Prototype, and Grill with Docs.
The opening workbench proposes these shared meanings:
- Clarify the decision before doing the work; proceed on reversible choices the user has delegated.
- Let the agent look up facts; reserve decisions for the human.
- Challenge overloaded terminology until one precise meaning is shared.
- Use a throwaway prototype only to settle a concrete uncertainty.
- Preserve a decision’s name and context as its identity.
These are proposed syntheses, not verbatim source quotations or demonstrated best practices. For example, the upstream wait-for-human stance and the delegated-progress interpretation are a real tension to disclose, not a difference to smooth away.
All five documents share one repository/author lineage. That makes them useful for demonstrating synthesis and recurrence, not independent corroboration from five unrelated sources.
Designed for editorial attention
The warm-paper visual language is functional: strong essence text, quiet metadata, fine dividing rules, and restrained olive actions. The important decision is visible before the machinery around it.


- Recorded before → after measurements:
- Desktop first essence: 420px → 249px from the top, about 41% earlier.
- Mobile first essence: 714px → 440px, about 38% earlier.
- Mobile table width: 928px → 343px at a 390px viewport.
- No horizontal page overflow in the exercised desktop or mobile states.
- Browser verification included:
- Empty-wording rejection, multiline editing, Escape, and Ctrl+Enter save.
- Search and filtered selection; review and split dialogs that fence stale state without discarding entered content.
- Evidence-only acceptance, restore, split/fuse/remove, guarded undo, and real file download.
- Replay deduplication and preservation of in-flight typing and focus.
- Two observed defects were repaired:
- Delayed dialog focus restoration stole fresh search input after reset; the repaired sequence passed eight consecutive repetitions.
- Long review dialogs opened at the bottom verdict control; mobile review now opens at scroll position zero with the comparison visible.
These are interface measurements and exercised scenarios—not a usability study, accessibility certification, or measured productivity gain. Test verdicts were disposable, not the user’s acceptance of the example skills.
07 · A deliberately local architecture
- Implemented source components:
- Svelte 5 and TypeScript client; Tailwind 4, daisyUI 5, and Lucide interface components.
- Local capability-checked HTTP API, server-sent state updates, and file-backed bench persistence.
- An OMP project extension/bridge for dispatch, checkpoints, and research results.
- Document identity, attributed impulses, essence proposals, revision records, and Markdown export.
- Separate demonstration surface:
- The editorial prototype is plain HTML, CSS, and JavaScript.
- No dependency installation, external font, CDN, or inference service is needed to open it.
- State is memory-only. Reload returns to the illustrative fixture.
- Hosting this pitch:
- A public, static Quartz companion wiki at
wiki.skll.loca.zone. - No application backend port and no application deployment.
- The demo attachment stays local to the browser; publishing it does not connect the production bridge.
- A public, static Quartz companion wiki at
Local capability checks are a boundary in the implementation, not a claim of a completed security audit or a production-hardened multi-user service.
08 · What is proven—and what is not
- Verified today, as recorded in the browser receipt:
- The editorial prototype and the specific scenarios described above.
- Genuine desktop/mobile captures, with no invented production screens.
- Implemented but not accepted end to end:
- The production client, local store/API, and OMP bridge source.
- Earlier isolated checks exist, but they do not prove the assembled live workflow.
- Required before production acceptance:
- Repair the outstanding production findings and align the approved prototype semantics.
- Perform the already-approved same-conversation restart to load the project extension.
- Correlate a real browser request with that conversation’s actual tool execution, durable evidence/proposals, and browser updates.
- Exercise interruption/recovery, identity handling, human review, and pinned export on that real path.
- Obtain the human’s actual acceptance. Neither green checks nor this page can substitute for it.
The outstanding production work, without euphemisms
The earlier read-only production audit identified ten unresolved issues. They were not fixed during the prototype or pitch rounds:
- Reactive startup can retrigger initialization.
- A wording-only acceptance guard blocks evidence-only review.
- Export preview invalidation misses corpus/selection changes.
- Structural mutations lack necessary revision checks.
- Saving can discard newer in-flight typing.
- Cross-request synthesis appends essences instead of reconciling shared meaning.
- Terminated research turns are not fully reconciled.
- An open pinned-evidence disclosure can go blank after acceptance.
- Known-mirror reconciliation is not exposed through the product.
- API recovery reasons are hidden from the interface.
Source review for this pitch also confirmed narrower production snapshots/restoration and different draft/context export semantics. Those are parity work, not capabilities to borrow from the prototype in a sales claim.
What this pitch deliberately does not claim
- More recurrence means higher confidence or independent corroboration.
- Preserved citations make a synthesized skill correct or validated.
- Recorded replay is live research.
- Existing components prove the entire system works.
- Prototype QA verdicts represent the user’s approval.
- A public wiki means the application is production-ready.
Try the idea. Inspect the evidence. Keep the decision yours.
The strongest next step is concrete: open the prototype, inspect one essence’s sources, change the wording, review an evidence-only revision, and compare the accepted export with the working state. Judge the actual editorial experience—not an abstract promise of autonomous intelligence.
Sources and receipts
- Prototype evidence:
- Public verification receipt: measured viewports, exercised behavior, limitations, artifact hashes, and the production boundary.
- The screenshots on this page are the preserved real-browser captures; the numbered SVGs are explanatory diagrams.
- The offline bundle contains both the editorial workbench and its original comparison prototype.
- Pinned upstream corpus:
- Wayfinder
- Grilling
- Domain Modeling
- Prototype
- Grill with Docs
- Pinned MIT license—Copyright © 2026 Matt Pocock. Retain upstream attribution; inspect the license before reuse.
- Technology references:
Evidence date: 7 September 2026. One page, one explainable journey: sources → meaning → explicit review → a skill you can take with you.

