Skip to content

Instantly share code, notes, and snippets.

@possibilities
Last active September 22, 2026 17:10
Show Gist options
  • Select an option

  • Save possibilities/dae32b35b4fc2d7a9b5cbcd92833bc32 to your computer and use it in GitHub Desktop.

Select an option

Save possibilities/dae32b35b4fc2d7a9b5cbcd92833bc32 to your computer and use it in GitHub Desktop.
DSPy + Jev toolbox doctrine — working doctrine for placing System One decisions and compiled LM programs in a software/agent stack
title DSPy + Jev toolbox doctrine
tags
agents
doctrine
dspy
jev
unfair-advantage
created 2026-09-22 16:47:53 UTC
updated 2026-09-22 17:10:00 UTC

DSPy + Jev toolbox doctrine

Status: working doctrine for our stack (rev 2 — incorporated external technical review 2026-09-22)
Stance: two sharp tools with different layers. Prefer precise placement. Combination only when the whole workflow improves. Use one, both, or neither.

Governing principle: Choose the runtime implementation empirically. Optimize behavior where success can be measured. Treat uncertainty as an input to policy, not permission to act. Keep authority in code. Require combinations to improve the complete workflow.

This document is split into facts, defaults, and invariants. Facts describe what the tools provide. Defaults say what to try first (measured exceptions allowed). Invariants are non-negotiable.


Facts

Jev

  • Hosted System One decision API (TypeSafe AI): send state + typed questions; get structured answers.
  • Primitives: choice (≤255 options), score (ordered 2–10 levels), noul (P(yes) ∈ [0,1]).
  • Closed output space: cannot invent options or emit invalid types. Can still pick the wrong valid option.
  • Role: a runtime model/API for bounded, typed judgments — not a text generator.

DSPy

  • Framework for programming and optimizing LM systems (Stanford NLP): signatures, modules, metrics, compile/promote.
  • Optimizers search instructions, demos, and sometimes weights (BootstrapFewShot, MIPROv2, GEPA, finetune paths).
  • Role: a framework for constructing and optimizing LM programs — classifiers, routers, extractors, generators, agents.
  • Compile is offline/batch; runtime runs a saved program until you recompile.

Layers, not mutually exclusive task categories

Jev and DSPy are different layers. A small answer shape does not imply a small reasoning task (“does this patch preserve behavior?” is yes/no but may need deep investigation). Prefer tools by fit and measurement, not by output type alone.

Probability semantics (do not conflate)

Returned value Meaning Consequence
Choice probabilities Distribution over the supplied options Relative to that candidate set
Choice/Score confidence Statistic from distribution shape Not automatically P(selected is correct)
Noul Estimated P(yes) Near 0 = strong no, not “low confidence”
Score Probability-weighted mean of level indices Not a percentage of a physical quantity

Define, per decision contract, which quantity drives policy and what acceptance / rejection / abstention / failure mean. Validate on representative data. Different decisions (notify vs block a dangerous command) need different thresholds.

Calibration is a batch property and a claim to verify on our data — not a guarantee about one answer, and it can degrade under distribution shift. Track both probability quality (e.g. reliability / Brier or log loss — carefully) and operational quality (error rate among auto-executed decisions, cost of those errors, automation coverage).


Defaults (try first; allow measured exceptions)

When to reach for which

Need Try first
Narrow bounded judgment on the hot path (route, gate, score, allow/deny, pick-among-known-N) Jev, after shortlist if N is large
Construct or optimize an LM program (including classifiers/routers) with a metric DSPy
Deterministic rules or a boring classifier already work Neither (or use as baseline)
Explanation / open-ended generation for humans LLM (often after a decision)

Smell tests

  • About to ask an LLM for JSON just to if on it → try Jev (or a simple baseline) first.
  • Quality is prompt folklore and success is measurable → DSPy is a candidate.
  • Bounded semantic decision → benchmark Jev vs simple baseline vs LM/DSPy; choose by quality, latency, cost, and operational fit — not by output type alone.
  • Keep Jev as a preferred candidate for narrow decisions, not their exclusive owner. TypeSafe notes limits on heavy reasoning indirection, numerical precision, and distracting context.

Softened former hard line

Don't use DSPy to pick among a known list.
Replace with: For bounded semantic decisions, prefer trying Jev (and a simple baseline). Use DSPy when building or optimizing an LM implementation is worthwhile. DSPy explicitly supports classification; excluding it a priori is not a technical limitation.

Canonical seams (use only when the workflow improves)

  1. Jev before LM — decide whether a generative / DSPy-compiled step should run.
  2. Jev beside LM — parallel risk/relevance checks; a required check must finish before the protected side effect (after-the-fact is monitoring, not prevention).
  3. Jev after LM — readiness/risk recommendation before send/apply/merge/pay — never as the sole authority.
  4. Jev routes among compiled programs — specialist DSPy modules vs one megaprompt.
  5. Choose vs create (workspaces, etc.) — see invariants + workspace flow below.
  6. Offline optimize Jev questions (hypothesis) — use measured outcomes + an optimizer (e.g. GEPA-style / custom adapter) to improve instructions/criteria that a cheap Jev call executes at runtime. Label as proposed custom integration, not a native DSPy–Jev feature.

Improvement loops

Component loops (attributable):

  • Jev: traces + outcomes → retune thresholds / rewrite criteria as situations → version question sets.
  • DSPy: traces + labels → metric → compile → held-out eval → promote or discard.

Shared end-to-end promotion test (required): component wins are not enough. A router can “improve” while sending work to the wrong downstream program; a gate can “improve” by escalating everything.

Label hygiene: do not harvest only uncertain cases — audit a sample of confident automations. Observed outcomes are noisy (merged ≠ correct; ignored notify ≠ unwanted; no complaint ≠ success).


Invariants (non-negotiable)

  1. Authority stays in code. Models may recommend and assess risk. Permissions, allowed targets, approvals, and execution constraints are enforced by code and downstream systems. A model result cannot grant permission the caller did not already have. Pattern: propose → deterministic authz → model risk/readiness → required approval → revalidate exact action + state → execute.
  2. Every consequential decision needs an explicit action policy (quantities, thresholds, abstain, failure) and an escalate/safe-default path. Mid probability is not a yes.
  3. Uncertainty is not evidence for creation. Prefer gather context / broaden search / human review before inventing new entities (workspaces, queues, labels).
  4. “None of the above” is an option you define when using Choice — plus confidence/fit checks. Relative best ≠ actually fits (consider best-candidate Choice + per-candidate noul/fit).
  5. Without credible evaluation, neither tool earns consequential autonomy. Start in suggestion/shadow mode; scale eval burden to consequence. Lack of labels is not a license to ship Jev (or DSPy) with real side effects.
  6. Pin and version model ids, question/criteria text, thresholds, state-building logic, DSPy program + dataset + metric + compile config. Add timeouts, bounded retries, stale-result handling, monitoring, and rollback — “escalate on uncertainty” does not cover outages.
  7. Combination requires a decision contract (below). No seam theater: the whole path (decision + retrieval + downstream + escalation + retries + amortized optimization + cost of mistakes) must beat the baseline.
  8. Provenance: distinguish documented API behavior, vendor-reported performance, and our hypotheses. Prefer TypeSafe’s own docs as authoritative for Jev API behavior; treat third-party aggregators as secondary. Date Jev-specific assumptions (early access announced ~2026-09-15; this doctrine rev 2026-09-22).

Workspace membership flow (revised)

Situations → next steps:

Situation Next step
One existing workspace fits clearly Assign (policy permitting)
Several plausibly fit Resolve ambiguity, or multi-assign if product allows
Supplied candidates don’t fit Broaden search before concluding none exists
Task lacks information Gather context / ask
No fit after adequate search Propose creation separately (policy, idempotent, concurrency-safe)
Model/API fails Explicit operational fallback

Notes:

  • Decide whether membership is exclusive; a single Choice silently encodes a product rule.
  • Measure retrieval separately from classification — omitted candidates cannot be recovered by confidence.
  • For small catalogs, try full catalog before shortlist complexity.
  • A DSPy (or LM) fallback that re-searches and chooses an existing workspace is legitimate — not only “create.”

Decision contract (required for each integration)

  • Decision: What question? What actions can follow?
  • Evidence: What state? Can candidates/facts be missing?
  • Baseline and challenger: What must the proposal beat?
  • Error policy: Cost of false positives vs false negatives?
  • Fallback: Ambiguity, missing info, timeout, outage?
  • Promotion test: Measured improvement + regression limits?
  • Release: Versions, owner, monitoring, rollback?

Anti-patterns

  • Frontier LLM for two-bit decisions on the hot path (without comparing to Jev/baseline)
  • DSPy compile with no metric / no promotion bar
  • Consequential Jev (or DSPy) autonomy without evaluation
  • Auto-action overnight on middling confidence / without code-enforced authz
  • Treating complementary nouls as logical proofs
  • Treating confidence as P(correct)
  • Creating entities because the model was uncertain
  • “We added both vendors” without end-to-end workflow improvement
  • Optimizing only on uncertain cases (misses confident mistakes)

Unfair advantage (what we’re chasing)

Precision of placement plus empirical honesty: generative cores that measurably improve, decision surfaces that stop wasting tokens on tiny judgments, authority that never leaves code, and promotion that survives an end-to-end test — with clear escalate paths and attributable component loops.


References / provenance (dated 2026-09-22)

  • DSPy: https://dspy.ai/ , https://github.com/stanfordnlp/dspy
  • TypeSafe Jev: prefer docs linked from TypeSafe’s own site / docs.typesafe.ai for API behavior
  • External review incorporated: ChatGPT share t_6ab2b5cb03048191a702220111a00967 (technical review of rev 1 doctrine + conversation gists)

Hand this to agents as-is. When a design argues for both tools, require a decision contract and name the seam — or justify a new one in one sentence tied to end-to-end improvement.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment