| title | DSPy + Jev toolbox doctrine | |||||
|---|---|---|---|---|---|---|
| tags |
|
|||||
| created | 2026-09-22 16:47:53 UTC | |||||
| updated | 2026-09-22 17:10:00 UTC |
Status: working doctrine for our stack (rev 2 — incorporated external technical review 2026-09-22)
Stance: two sharp tools with different layers. Prefer precise placement. Combination only when the whole workflow improves. Use one, both, or neither.
Governing principle: Choose the runtime implementation empirically. Optimize behavior where success can be measured. Treat uncertainty as an input to policy, not permission to act. Keep authority in code. Require combinations to improve the complete workflow.
This document is split into facts, defaults, and invariants. Facts describe what the tools provide. Defaults say what to try first (measured exceptions allowed). Invariants are non-negotiable.
- Hosted System One decision API (TypeSafe AI): send
state+ typedquestions; get structuredanswers. - Primitives: choice (≤255 options), score (ordered 2–10 levels), noul (P(yes) ∈ [0,1]).
- Closed output space: cannot invent options or emit invalid types. Can still pick the wrong valid option.
- Role: a runtime model/API for bounded, typed judgments — not a text generator.
- Framework for programming and optimizing LM systems (Stanford NLP): signatures, modules, metrics, compile/promote.
- Optimizers search instructions, demos, and sometimes weights (BootstrapFewShot, MIPROv2, GEPA, finetune paths).
- Role: a framework for constructing and optimizing LM programs — classifiers, routers, extractors, generators, agents.
- Compile is offline/batch; runtime runs a saved program until you recompile.
Jev and DSPy are different layers. A small answer shape does not imply a small reasoning task (“does this patch preserve behavior?” is yes/no but may need deep investigation). Prefer tools by fit and measurement, not by output type alone.
| Returned value | Meaning | Consequence |
|---|---|---|
| Choice probabilities | Distribution over the supplied options | Relative to that candidate set |
| Choice/Score confidence | Statistic from distribution shape | Not automatically P(selected is correct) |
| Noul | Estimated P(yes) | Near 0 = strong no, not “low confidence” |
| Score | Probability-weighted mean of level indices | Not a percentage of a physical quantity |
Define, per decision contract, which quantity drives policy and what acceptance / rejection / abstention / failure mean. Validate on representative data. Different decisions (notify vs block a dangerous command) need different thresholds.
Calibration is a batch property and a claim to verify on our data — not a guarantee about one answer, and it can degrade under distribution shift. Track both probability quality (e.g. reliability / Brier or log loss — carefully) and operational quality (error rate among auto-executed decisions, cost of those errors, automation coverage).
| Need | Try first |
|---|---|
| Narrow bounded judgment on the hot path (route, gate, score, allow/deny, pick-among-known-N) | Jev, after shortlist if N is large |
| Construct or optimize an LM program (including classifiers/routers) with a metric | DSPy |
| Deterministic rules or a boring classifier already work | Neither (or use as baseline) |
| Explanation / open-ended generation for humans | LLM (often after a decision) |
Smell tests
- About to ask an LLM for JSON just to
ifon it → try Jev (or a simple baseline) first. - Quality is prompt folklore and success is measurable → DSPy is a candidate.
- Bounded semantic decision → benchmark Jev vs simple baseline vs LM/DSPy; choose by quality, latency, cost, and operational fit — not by output type alone.
- Keep Jev as a preferred candidate for narrow decisions, not their exclusive owner. TypeSafe notes limits on heavy reasoning indirection, numerical precision, and distracting context.
Don't use DSPy to pick among a known list.
Replace with: For bounded semantic decisions, prefer trying Jev (and a simple baseline). Use DSPy when building or optimizing an LM implementation is worthwhile. DSPy explicitly supports classification; excluding it a priori is not a technical limitation.
- Jev before LM — decide whether a generative / DSPy-compiled step should run.
- Jev beside LM — parallel risk/relevance checks; a required check must finish before the protected side effect (after-the-fact is monitoring, not prevention).
- Jev after LM — readiness/risk recommendation before send/apply/merge/pay — never as the sole authority.
- Jev routes among compiled programs — specialist DSPy modules vs one megaprompt.
- Choose vs create (workspaces, etc.) — see invariants + workspace flow below.
- Offline optimize Jev questions (hypothesis) — use measured outcomes + an optimizer (e.g. GEPA-style / custom adapter) to improve instructions/criteria that a cheap Jev call executes at runtime. Label as proposed custom integration, not a native DSPy–Jev feature.
Component loops (attributable):
- Jev: traces + outcomes → retune thresholds / rewrite criteria as situations → version question sets.
- DSPy: traces + labels → metric → compile → held-out eval → promote or discard.
Shared end-to-end promotion test (required): component wins are not enough. A router can “improve” while sending work to the wrong downstream program; a gate can “improve” by escalating everything.
Label hygiene: do not harvest only uncertain cases — audit a sample of confident automations. Observed outcomes are noisy (merged ≠ correct; ignored notify ≠ unwanted; no complaint ≠ success).
- Authority stays in code. Models may recommend and assess risk. Permissions, allowed targets, approvals, and execution constraints are enforced by code and downstream systems. A model result cannot grant permission the caller did not already have. Pattern: propose → deterministic authz → model risk/readiness → required approval → revalidate exact action + state → execute.
- Every consequential decision needs an explicit action policy (quantities, thresholds, abstain, failure) and an escalate/safe-default path. Mid probability is not a yes.
- Uncertainty is not evidence for creation. Prefer gather context / broaden search / human review before inventing new entities (workspaces, queues, labels).
- “None of the above” is an option you define when using Choice — plus confidence/fit checks. Relative best ≠ actually fits (consider best-candidate Choice + per-candidate noul/fit).
- Without credible evaluation, neither tool earns consequential autonomy. Start in suggestion/shadow mode; scale eval burden to consequence. Lack of labels is not a license to ship Jev (or DSPy) with real side effects.
- Pin and version model ids, question/criteria text, thresholds, state-building logic, DSPy program + dataset + metric + compile config. Add timeouts, bounded retries, stale-result handling, monitoring, and rollback — “escalate on uncertainty” does not cover outages.
- Combination requires a decision contract (below). No seam theater: the whole path (decision + retrieval + downstream + escalation + retries + amortized optimization + cost of mistakes) must beat the baseline.
- Provenance: distinguish documented API behavior, vendor-reported performance, and our hypotheses. Prefer TypeSafe’s own docs as authoritative for Jev API behavior; treat third-party aggregators as secondary. Date Jev-specific assumptions (early access announced ~2026-09-15; this doctrine rev 2026-09-22).
Situations → next steps:
| Situation | Next step |
|---|---|
| One existing workspace fits clearly | Assign (policy permitting) |
| Several plausibly fit | Resolve ambiguity, or multi-assign if product allows |
| Supplied candidates don’t fit | Broaden search before concluding none exists |
| Task lacks information | Gather context / ask |
| No fit after adequate search | Propose creation separately (policy, idempotent, concurrency-safe) |
| Model/API fails | Explicit operational fallback |
Notes:
- Decide whether membership is exclusive; a single Choice silently encodes a product rule.
- Measure retrieval separately from classification — omitted candidates cannot be recovered by confidence.
- For small catalogs, try full catalog before shortlist complexity.
- A DSPy (or LM) fallback that re-searches and chooses an existing workspace is legitimate — not only “create.”
- Decision: What question? What actions can follow?
- Evidence: What state? Can candidates/facts be missing?
- Baseline and challenger: What must the proposal beat?
- Error policy: Cost of false positives vs false negatives?
- Fallback: Ambiguity, missing info, timeout, outage?
- Promotion test: Measured improvement + regression limits?
- Release: Versions, owner, monitoring, rollback?
- Frontier LLM for two-bit decisions on the hot path (without comparing to Jev/baseline)
- DSPy compile with no metric / no promotion bar
- Consequential Jev (or DSPy) autonomy without evaluation
- Auto-action overnight on middling confidence / without code-enforced authz
- Treating complementary nouls as logical proofs
- Treating confidence as P(correct)
- Creating entities because the model was uncertain
- “We added both vendors” without end-to-end workflow improvement
- Optimizing only on uncertain cases (misses confident mistakes)
Precision of placement plus empirical honesty: generative cores that measurably improve, decision surfaces that stop wasting tokens on tiny judgments, authority that never leaves code, and promotion that survives an end-to-end test — with clear escalate paths and attributable component loops.
- DSPy: https://dspy.ai/ , https://github.com/stanfordnlp/dspy
- TypeSafe Jev: prefer docs linked from TypeSafe’s own site / docs.typesafe.ai for API behavior
- External review incorporated: ChatGPT share
t_6ab2b5cb03048191a702220111a00967(technical review of rev 1 doctrine + conversation gists)
Hand this to agents as-is. When a design argues for both tools, require a decision contract and name the seam — or justify a new one in one sentence tied to end-to-end improvement.