Skip to content

Instantly share code, notes, and snippets.

Show Gist options
  • Select an option

  • Save hammer/956357c87da62bd41a3b8e323caecf7b to your computer and use it in GitHub Desktop.

Select an option

Save hammer/956357c87da62bd41a3b8e323caecf7b to your computer and use it in GitHub Desktop.
Covering the competency map and the task curriculum: where Marin's RL environments come from
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Covering the competency map and the task curriculum: where Marin&#x27;s RL environments come from</title>
<style>
:root {
--ink: #1a1a2e; --ink-secondary: #555770; --ink-faint: #8a8a9a;
--accent: #8b2500; --accent-light: #c4530a;
--surface: #faf9f6; --surface-raised: #f0eeea; --surface-code: #f5f3ef;
--rule: #d4d0c8; --link: #8b2500; --link-hover: #c4530a; --ref-bg: #f7f5f1;
--content-width: 650px; --sidenote-width: 230px; --sidenote-gap: 30px;
}
@media (prefers-color-scheme: dark) {
:root:not([data-theme="light"]) {
--ink: #d8d5cf; --ink-secondary: #9e9bab; --ink-faint: #6e6b7b;
--accent: #d4764e; --accent-light: #e8956e;
--surface: #1a1a24; --surface-raised: #242430; --surface-code: #20202c;
--rule: #33333f; --link: #d4764e; --link-hover: #e8956e; --ref-bg: #1e1e2a;
}
}
:root[data-theme="dark"] {
--ink: #d8d5cf; --ink-secondary: #9e9bab; --ink-faint: #6e6b7b;
--accent: #d4764e; --accent-light: #e8956e;
--surface: #1a1a24; --surface-raised: #242430; --surface-code: #20202c;
--rule: #33333f; --link: #d4764e; --link-hover: #e8956e; --ref-bg: #1e1e2a;
}
* { margin: 0; padding: 0; box-sizing: border-box; }
body {
background: var(--surface); color: var(--ink);
font-family: 'Helvetica Neue', Helvetica, Arial, sans-serif;
font-size: 16px; line-height: 1.7;
-webkit-font-smoothing: antialiased;
}
.page {
max-width: calc(var(--content-width) + var(--sidenote-width) + var(--sidenote-gap) + 80px);
margin: 0 auto; padding: 3rem 40px 4rem; position: relative;
}
@media (max-width: 1060px) {
.page { max-width: 100%; padding: 2rem 1.5rem 3rem; }
}
.content { max-width: var(--content-width); }
.paper-header { max-width: var(--content-width); margin-bottom: 2.5rem; padding-bottom: 2rem; }
.paper-header h1 {
font-family: 'Palatino Linotype', Palatino, 'Book Antiqua', Georgia, serif;
font-size: 2rem; font-weight: 600; line-height: 1.25; color: var(--ink);
text-wrap: balance; margin-bottom: 0.6rem; letter-spacing: -0.01em;
}
.paper-meta { font-size: 0.875rem; color: var(--ink-secondary); line-height: 1.5; }
.paper-meta .author { font-weight: 500; }
.paper-meta .mumwelt-link { color: var(--ink-secondary); text-decoration: none; border-bottom: 1px dotted var(--ink-faint); }
.paper-meta .mumwelt-link:hover { color: var(--accent); border-bottom-color: var(--accent); }
.prompt-box {
max-width: var(--content-width); margin-bottom: 2.5rem; position: relative;
}
.prompt-label {
position: absolute; left: -0.6em; top: 50%; transform: translateY(-50%);
font-family: 'Palatino Linotype', Palatino, 'Book Antiqua', Georgia, serif;
font-size: 5rem; font-weight: 700; color: var(--ink); opacity: 0.08;
line-height: 1; pointer-events: none; user-select: none;
}
.prompt-text {
font-family: 'Palatino Linotype', Palatino, 'Book Antiqua', Georgia, serif;
font-style: italic; font-size: 1.15rem; line-height: 1.55; color: var(--ink);
position: relative;
}
.abstract { margin-bottom: 2.5rem; max-width: var(--content-width); }
.abstract-label {
font-size: 0.7rem; font-weight: 600; text-transform: uppercase;
letter-spacing: 0.1em; color: var(--ink-secondary); margin-bottom: 0.5rem;
}
.abstract p { font-size: 0.92rem; line-height: 1.75; color: var(--ink); }
h1, h2, h3 {
font-family: 'Palatino Linotype', Palatino, 'Book Antiqua', Georgia, serif;
font-weight: 600; color: var(--ink); text-wrap: balance;
}
h2 { font-size: 1.4rem; margin-top: 2.5rem; margin-bottom: 0.75rem; letter-spacing: -0.005em; }
h3 { font-size: 1.1rem; margin-top: 1.75rem; margin-bottom: 0.5rem; }
p { margin-bottom: 1rem; max-width: var(--content-width); }
a { color: var(--link); text-decoration: none; border-bottom: 1px solid transparent;
transition: border-color 0.15s, color 0.15s; }
a:hover { color: var(--link-hover); border-bottom-color: var(--link-hover); }
.date-label { cursor: default; border-bottom: 1px dotted var(--ink-faint); }
strong { font-weight: 600; }
.sidenote-checkbox { display: none; }
.sidenote-toggle { display: none; }
.sidenote {
float: right; clear: right; width: var(--sidenote-width);
margin-right: calc(-1 * (var(--sidenote-width) + var(--sidenote-gap)));
margin-top: 0.2rem; margin-bottom: 1rem;
font-size: 0.8rem; line-height: 1.5; color: var(--ink-secondary);
}
.sidenote-number { font-size: 0.7rem; font-weight: 600; color: var(--accent); margin-right: 0.3em; }
@media (max-width: 1060px) {
.sidenote-toggle {
display: inline; cursor: pointer; color: var(--accent);
font-size: 0.78rem; font-weight: 600; user-select: none;
}
.sidenote {
float: none; display: none; width: 100%; margin: 0.4rem 0 0.75rem 0;
font-size: 0.84rem; padding: 0.6rem 0.9rem; background: var(--surface-raised);
border-radius: 4px; border-left: 2px solid var(--accent);
}
.sidenote-checkbox:checked + .sidenote { display: block; }
.sidenote-number { display: none; }
}
blockquote { border-left: 2px solid var(--rule); padding-left: 1.25rem;
margin: 1.25rem 0; color: var(--ink-secondary); font-style: italic; }
code { font-family: 'SF Mono', Menlo, Consolas, monospace; font-size: 0.85em;
background: var(--surface-code); padding: 0.15em 0.35em; border-radius: 3px; }
pre { background: var(--surface-code); border: 1px solid var(--rule); border-radius: 4px;
padding: 1rem 1.25rem; overflow-x: auto; margin: 1.25rem 0; max-width: var(--content-width); }
pre code { background: none; padding: 0; font-size: 0.82rem; line-height: 1.6; }
ul, ol { margin-bottom: 1rem; padding-left: 1.5rem; max-width: var(--content-width); }
li { margin-bottom: 0.35rem; }
li::marker { color: var(--ink-faint); }
table { max-width: var(--content-width); border-collapse: collapse; width: 100%;
margin: 1.25rem 0; font-variant-numeric: tabular-nums; font-size: 0.9rem; }
thead { border-top: 2px solid var(--ink); border-bottom: 1px solid var(--ink); }
th { font-weight: 600; padding: 0.3rem 0.75rem 0.35rem; text-align: left;
line-height: 1.2; white-space: nowrap; font-size: 0.82rem; vertical-align: bottom; }
td { padding: 0.35rem 0.75rem; border: none; vertical-align: top; }
tbody { border-bottom: 1.5px solid var(--ink); }
th:first-child, td:first-child { padding-left: 0; }
th:last-child, td:last-child { padding-right: 0; }
figure { margin: 2rem 0; max-width: var(--content-width); }
figure img { width: 100%; border-radius: 3px; border: 1px solid var(--rule); }
figcaption { font-size: 0.8rem; color: var(--ink-secondary); margin-top: 0.5rem;
line-height: 1.5; font-style: italic; }
footer.provenance {
margin-top: 3rem; padding-top: 1rem; max-width: var(--content-width);
font-size: 0.78rem; line-height: 1.6; color: var(--ink-faint);
}
footer.provenance p { margin-bottom: 0.3rem; }
footer.provenance a { color: var(--ink-faint); }
footer.provenance blockquote { margin: 0; padding: 0; border: none; color: inherit; }
a[data-hover-title] { position: relative; }
.hover-card {
position: absolute; bottom: 100%; left: 50%; transform: translateX(-50%);
width: 320px; max-width: 90vw; padding: 0.65rem 0.8rem;
background: var(--surface-raised); border: 1px solid var(--rule);
border-radius: 6px; box-shadow: 0 4px 12px rgba(0,0,0,0.1);
font-size: 0.78rem; line-height: 1.45; color: var(--ink);
pointer-events: none; z-index: 100; margin-bottom: 6px;
opacity: 0; transition: opacity 0.12s;
}
a[data-hover-title]:hover .hover-card,
a[data-hover-title]:focus .hover-card,
a.cite:hover .hover-card { opacity: 1; }
.hover-card .hc-title { font-weight: 600; margin-bottom: 0.2rem; }
.hover-card .hc-meta { font-size: 0.72rem; color: var(--ink-faint); margin-bottom: 0.25rem; }
.hover-card .hc-status {
display: inline-block; font-size: 0.65rem; font-weight: 600;
text-transform: uppercase; letter-spacing: 0.04em;
padding: 0.1em 0.4em; border-radius: 3px; margin-right: 0.4em;
}
.hc-status-open { background: #fff3e0; color: #e65100; }
.hc-status-closed, .hc-status-merged { background: #e8f5e9; color: #2e7d32; }
@media (prefers-color-scheme: dark) {
:root:not([data-theme="light"]) .hover-card { box-shadow: 0 4px 12px rgba(0,0,0,0.3); }
:root:not([data-theme="light"]) .hc-status-open { background: #3a2a10; color: #ffb74d; }
:root:not([data-theme="light"]) .hc-status-closed,
:root:not([data-theme="light"]) .hc-status-merged { background: #1b3a1e; color: #66bb6a; }
}
:root[data-theme="dark"] .hover-card { box-shadow: 0 4px 12px rgba(0,0,0,0.3); }
:root[data-theme="dark"] .hc-status-open { background: #3a2a10; color: #ffb74d; }
:root[data-theme="dark"] .hc-status-closed,
:root[data-theme="dark"] .hc-status-merged { background: #1b3a1e; color: #66bb6a; }
.hover-card .hc-desc { color: var(--ink-secondary); }
.cite {
font-size: 0.72rem; vertical-align: super; line-height: 0;
color: var(--accent); font-weight: 600; text-decoration: none;
border-bottom: none !important; position: relative;
}
.cite:hover { color: var(--link-hover); }
.references { margin-top: 3rem; max-width: var(--content-width); }
.references h2 { font-size: 1.15rem; margin-bottom: 1rem; }
.ref-list { list-style: none; padding: 0; counter-reset: ref; }
.ref-list li {
counter-increment: ref; display: flex; align-items: baseline;
gap: 0.5em; font-size: 0.82rem; line-height: 1.55;
margin-bottom: 0.4rem; color: var(--ink-secondary);
}
.ref-list li::before {
content: "[" counter(ref) "]"; flex-shrink: 0;
font-variant-numeric: tabular-nums; color: var(--ink-faint);
font-size: 0.78rem; min-width: 2.2em;
}
.ref-list .ref-body { flex: 1; min-width: 0; }
.ref-list .ref-title { font-weight: 500; color: var(--ink); }
.ref-list .ref-url {
font-family: 'SF Mono', Menlo, Consolas, monospace; font-size: 0.75rem;
color: var(--ink-faint); word-break: break-all; margin-left: 0.4em;
}
.ref-list .ref-url a { color: var(--ink-faint); border-bottom: none; }
.ref-list .ref-url a:hover { color: var(--link-hover); }
.ref-back {
color: var(--accent); text-decoration: none; border-bottom: none !important;
margin-left: 0.3em; font-size: 0.78rem;
}
.ref-back:hover { color: var(--link-hover); }
.status {
display: inline-block; font-size: 0.65rem; font-weight: 600;
text-transform: uppercase; letter-spacing: 0.04em;
padding: 0.15em 0.5em; border-radius: 3px; vertical-align: middle;
}
.status-done { background: #e8f5e9; color: #2e7d32; }
.status-open { background: #fff3e0; color: #e65100; }
.status-blocked { background: #fce4ec; color: #c62828; }
@media (prefers-color-scheme: dark) {
:root:not([data-theme="light"]) .status-done { background: #1b3a1e; color: #66bb6a; }
:root:not([data-theme="light"]) .status-open { background: #3a2a10; color: #ffb74d; }
:root:not([data-theme="light"]) .status-blocked { background: #3a1520; color: #ef9a9a; }
}
:root[data-theme="dark"] .status-done { background: #1b3a1e; color: #66bb6a; }
:root[data-theme="dark"] .status-open { background: #3a2a10; color: #ffb74d; }
:root[data-theme="dark"] .status-blocked { background: #3a1520; color: #ef9a9a; }
.cite-card {
position: absolute; bottom: 100%; left: 50%; transform: translateX(-50%);
width: 360px; max-width: 90vw; padding: 0.65rem 0.8rem;
background: var(--surface-raised); border: 1px solid var(--rule);
border-radius: 6px; box-shadow: 0 4px 12px rgba(0,0,0,0.1);
font-size: 0.78rem; line-height: 1.45; color: var(--ink);
pointer-events: none; z-index: 100; margin-bottom: 6px;
opacity: 0; transition: opacity 0.12s;
font-weight: 400; vertical-align: baseline; text-align: left;
}
a.cite:hover .cite-card { opacity: 1; }
@media (prefers-color-scheme: dark) {
:root:not([data-theme="light"]) .cite-card { box-shadow: 0 4px 12px rgba(0,0,0,0.3); }
}
:root[data-theme="dark"] .cite-card { box-shadow: 0 4px 12px rgba(0,0,0,0.3); }
@media print {
body { font-size: 11pt; }
.page { max-width: 100%; padding: 0; }
.sidenote { float: right; width: 180px; margin-right: -210px; }
.sidenote-toggle { display: none; }
a { color: inherit; border-bottom: none; }
.hover-card { display: none; }
.cite-card { display: none; }
}
.katex{font:normal 1.21em KaTeX_Main,Times New Roman,serif;line-height:1.2;position:relative;text-indent:0;text-rendering:auto}.katex *{-ms-high-contrast-adjust:none!important;border-color:currentColor}.katex .katex-version:after{content:"0.18.7"}.katex .katex-mathml{border:0;-webkit-clip-path:inset(50%);clip-path:inset(50%);height:1px;overflow:hidden;padding:0;position:absolute;width:1px}.katex .katex-html>.katex-newline{display:block}.katex .katex-base{position:relative;white-space:nowrap;width:-webkit-min-content;width:-moz-min-content;width:min-content}.katex .katex-base,.katex .katex-strut{display:inline-block}.katex .textbf{font-weight:700}.katex .textit{font-style:italic}.katex .textrm{font-family:KaTeX_Main}.katex .textsf{font-family:KaTeX_SansSerif}.katex .texttt{font-family:KaTeX_Typewriter}.katex .mathnormal{font-family:KaTeX_Math;font-style:italic}.katex .mathit{font-family:KaTeX_Main;font-style:italic}.katex .mathrm{font-style:normal}.katex .mathbf{font-family:KaTeX_Main;font-weight:700}.katex .boldsymbol{font-family:KaTeX_Math;font-style:italic;font-weight:700}.katex .amsrm,.katex .mathbb,.katex .textbb{font-family:KaTeX_AMS}.katex .mathcal{font-family:KaTeX_Caligraphic}.katex .mathfrak,.katex .textfrak{font-family:KaTeX_Fraktur}.katex .mathboldfrak,.katex .textboldfrak{font-family:KaTeX_Fraktur;font-weight:700}.katex .mathtt{font-family:KaTeX_Typewriter}.katex .mathscr,.katex .textscr{font-family:KaTeX_Script}.katex .mathsf,.katex .textsf{font-family:KaTeX_SansSerif}.katex .mathboldsf,.katex .textboldsf{font-family:KaTeX_SansSerif;font-weight:700}.katex .mathitsf,.katex .mathsfit,.katex .textitsf{font-family:KaTeX_SansSerif;font-style:italic}.katex .mainrm{font-family:KaTeX_Main;font-style:normal}.katex .vlist-t{border-collapse:collapse;display:inline-table;table-layout:fixed}.katex .vlist-r{display:table-row}.katex .vlist{display:table-cell;position:relative;vertical-align:bottom}.katex .vlist>span{display:block;height:0;position:relative}.katex .vlist>span>span{display:inline-block}.katex .vlist>span>.pstrut{overflow:hidden;width:0}.katex .vlist-t2{margin-right:-2px}.katex .vlist-s{display:table-cell;font-size:1px;min-width:2px;vertical-align:bottom;width:2px}.katex .katex-vbox{align-items:baseline;display:inline-flex;flex-direction:column}.katex .katex-thinbox{display:inline-flex;flex-direction:row;max-width:0;width:0}.katex .msupsub{text-align:left}.katex .mfrac>span>span{text-align:center}.katex .mfrac .frac-line{border-bottom-style:solid;display:inline-block;width:100%}.katex .katex-hdashline,.katex .katex-hline,.katex .katex-overline .overline-line,.katex .katex-rule,.katex .katex-underline .underline-line,.katex .mfrac .frac-line{min-height:1px}.katex .mspace{display:inline-block}.katex .katex-smash{display:inline;line-height:0}.katex .clap,.katex .llap,.katex .rlap{position:relative;width:0}.katex .clap>.katex-inner,.katex .llap>.katex-inner,.katex .rlap>.katex-inner{position:absolute}.katex .clap>.katex-fix,.katex .llap>.katex-fix,.katex .rlap>.katex-fix{display:inline-block}.katex .llap>.katex-inner{right:0}.katex .clap>.katex-inner,.katex .rlap>.katex-inner{left:0}.katex .clap>.katex-inner>span{margin-left:-50%;margin-right:50%}.katex .katex-rule{border:0 solid;display:inline-block;position:relative}.katex .katex-hline,.katex .katex-overline .overline-line,.katex .katex-underline .underline-line{border-bottom-style:solid;display:inline-block;width:100%}.katex .katex-hdashline{border-bottom-style:dashed;display:inline-block;width:100%}.katex .sqrt>.katex-root{margin-left:.2777777778em;margin-right:-.5555555556em}.katex .fontsize-ensurer.reset-size1.size1,.katex .katex-sizing.reset-size1.size1{font-size:1em}.katex .fontsize-ensurer.reset-size1.size2,.katex .katex-sizing.reset-size1.size2{font-size:1.2em}.katex .fontsize-ensurer.reset-size1.size3,.katex .katex-sizing.reset-size1.size3{font-size:1.4em}.katex .fontsize-ensurer.reset-size1.size4,.katex .katex-sizing.reset-size1.size4{font-size:1.6em}.katex .fontsize-ensurer.reset-size1.size5,.katex .katex-sizing.reset-size1.size5{font-size:1.8em}.katex .fontsize-ensurer.reset-size1.size6,.katex .katex-sizing.reset-size1.size6{font-size:2em}.katex .fontsize-ensurer.reset-size1.size7,.katex .katex-sizing.reset-size1.size7{font-size:2.4em}.katex .fontsize-ensurer.reset-size1.size8,.katex .katex-sizing.reset-size1.size8{font-size:2.88em}.katex .fontsize-ensurer.reset-size1.size9,.katex .katex-sizing.reset-size1.size9{font-size:3.456em}.katex .fontsize-ensurer.reset-size1.size10,.katex .katex-sizing.reset-size1.size10{font-size:4.148em}.katex .fontsize-ensurer.reset-size1.size11,.katex .katex-sizing.reset-size1.size11{font-size:4.976em}.katex .fontsize-ensurer.reset-size2.size1,.katex .katex-sizing.reset-size2.size1{font-size:.8333333333em}.katex .fontsize-ensurer.reset-size2.size2,.katex .katex-sizing.reset-size2.size2{font-size:1em}.katex .fontsize-ensurer.reset-size2.size3,.katex .katex-sizing.reset-size2.size3{font-size:1.1666666667em}.katex .fontsize-ensurer.reset-size2.size4,.katex .katex-sizing.reset-size2.size4{font-size:1.3333333333em}.katex .fontsize-ensurer.reset-size2.size5,.katex .katex-sizing.reset-size2.size5{font-size:1.5em}.katex .fontsize-ensurer.reset-size2.size6,.katex .katex-sizing.reset-size2.size6{font-size:1.6666666667em}.katex .fontsize-ensurer.reset-size2.size7,.katex .katex-sizing.reset-size2.size7{font-size:2em}.katex .fontsize-ensurer.reset-size2.size8,.katex .katex-sizing.reset-size2.size8{font-size:2.4em}.katex .fontsize-ensurer.reset-size2.size9,.katex .katex-sizing.reset-size2.size9{font-size:2.88em}.katex .fontsize-ensurer.reset-size2.size10,.katex .katex-sizing.reset-size2.size10{font-size:3.4566666667em}.katex .fontsize-ensurer.reset-size2.size11,.katex .katex-sizing.reset-size2.size11{font-size:4.1466666667em}.katex .fontsize-ensurer.reset-size3.size1,.katex .katex-sizing.reset-size3.size1{font-size:.7142857143em}.katex .fontsize-ensurer.reset-size3.size2,.katex .katex-sizing.reset-size3.size2{font-size:.8571428571em}.katex .fontsize-ensurer.reset-size3.size3,.katex .katex-sizing.reset-size3.size3{font-size:1em}.katex .fontsize-ensurer.reset-size3.size4,.katex .katex-sizing.reset-size3.size4{font-size:1.1428571429em}.katex .fontsize-ensurer.reset-size3.size5,.katex .katex-sizing.reset-size3.size5{font-size:1.2857142857em}.katex .fontsize-ensurer.reset-size3.size6,.katex .katex-sizing.reset-size3.size6{font-size:1.4285714286em}.katex .fontsize-ensurer.reset-size3.size7,.katex .katex-sizing.reset-size3.size7{font-size:1.7142857143em}.katex .fontsize-ensurer.reset-size3.size8,.katex .katex-sizing.reset-size3.size8{font-size:2.0571428571em}.katex .fontsize-ensurer.reset-size3.size9,.katex .katex-sizing.reset-size3.size9{font-size:2.4685714286em}.katex .fontsize-ensurer.reset-size3.size10,.katex .katex-sizing.reset-size3.size10{font-size:2.9628571429em}.katex .fontsize-ensurer.reset-size3.size11,.katex .katex-sizing.reset-size3.size11{font-size:3.5542857143em}.katex .fontsize-ensurer.reset-size4.size1,.katex .katex-sizing.reset-size4.size1{font-size:.625em}.katex .fontsize-ensurer.reset-size4.size2,.katex .katex-sizing.reset-size4.size2{font-size:.75em}.katex .fontsize-ensurer.reset-size4.size3,.katex .katex-sizing.reset-size4.size3{font-size:.875em}.katex .fontsize-ensurer.reset-size4.size4,.katex .katex-sizing.reset-size4.size4{font-size:1em}.katex .fontsize-ensurer.reset-size4.size5,.katex .katex-sizing.reset-size4.size5{font-size:1.125em}.katex .fontsize-ensurer.reset-size4.size6,.katex .katex-sizing.reset-size4.size6{font-size:1.25em}.katex .fontsize-ensurer.reset-size4.size7,.katex .katex-sizing.reset-size4.size7{font-size:1.5em}.katex .fontsize-ensurer.reset-size4.size8,.katex .katex-sizing.reset-size4.size8{font-size:1.8em}.katex .fontsize-ensurer.reset-size4.size9,.katex .katex-sizing.reset-size4.size9{font-size:2.16em}.katex .fontsize-ensurer.reset-size4.size10,.katex .katex-sizing.reset-size4.size10{font-size:2.5925em}.katex .fontsize-ensurer.reset-size4.size11,.katex .katex-sizing.reset-size4.size11{font-size:3.11em}.katex .fontsize-ensurer.reset-size5.size1,.katex .katex-sizing.reset-size5.size1{font-size:.5555555556em}.katex .fontsize-ensurer.reset-size5.size2,.katex .katex-sizing.reset-size5.size2{font-size:.6666666667em}.katex .fontsize-ensurer.reset-size5.size3,.katex .katex-sizing.reset-size5.size3{font-size:.7777777778em}.katex .fontsize-ensurer.reset-size5.size4,.katex .katex-sizing.reset-size5.size4{font-size:.8888888889em}.katex .fontsize-ensurer.reset-size5.size5,.katex .katex-sizing.reset-size5.size5{font-size:1em}.katex .fontsize-ensurer.reset-size5.size6,.katex .katex-sizing.reset-size5.size6{font-size:1.1111111111em}.katex .fontsize-ensurer.reset-size5.size7,.katex .katex-sizing.reset-size5.size7{font-size:1.3333333333em}.katex .fontsize-ensurer.reset-size5.size8,.katex .katex-sizing.reset-size5.size8{font-size:1.6em}.katex .fontsize-ensurer.reset-size5.size9,.katex .katex-sizing.reset-size5.size9{font-size:1.92em}.katex .fontsize-ensurer.reset-size5.size10,.katex .katex-sizing.reset-size5.size10{font-size:2.3044444444em}.katex .fontsize-ensurer.reset-size5.size11,.katex .katex-sizing.reset-size5.size11{font-size:2.7644444444em}.katex .fontsize-ensurer.reset-size6.size1,.katex .katex-sizing.reset-size6.size1{font-size:.5em}.katex .fontsize-ensurer.reset-size6.size2,.katex .katex-sizing.reset-size6.size2{font-size:.6em}.katex .fontsize-ensurer.reset-size6.size3,.katex .katex-sizing.reset-size6.size3{font-size:.7em}.katex .fontsize-ensurer.reset-size6.size4,.katex .katex-sizing.reset-size6.size4{font-size:.8em}.katex .fontsize-ensurer.reset-size6.size5,.katex .katex-sizing.reset-size6.size5{font-size:.9em}.katex .fontsize-ensurer.reset-size6.size6,.katex .katex-sizing.reset-size6.size6{font-size:1em}.katex .fontsize-ensurer.reset-size6.size7,.katex .katex-sizing.reset-size6.size7{font-size:1.2em}.katex .fontsize-ensurer.reset-size6.size8,.katex .katex-sizing.reset-size6.size8{font-size:1.44em}.katex .fontsize-ensurer.reset-size6.size9,.katex .katex-sizing.reset-size6.size9{font-size:1.728em}.katex .fontsize-ensurer.reset-size6.size10,.katex .katex-sizing.reset-size6.size10{font-size:2.074em}.katex .fontsize-ensurer.reset-size6.size11,.katex .katex-sizing.reset-size6.size11{font-size:2.488em}.katex .fontsize-ensurer.reset-size7.size1,.katex .katex-sizing.reset-size7.size1{font-size:.4166666667em}.katex .fontsize-ensurer.reset-size7.size2,.katex .katex-sizing.reset-size7.size2{font-size:.5em}.katex .fontsize-ensurer.reset-size7.size3,.katex .katex-sizing.reset-size7.size3{font-size:.5833333333em}.katex .fontsize-ensurer.reset-size7.size4,.katex .katex-sizing.reset-size7.size4{font-size:.6666666667em}.katex .fontsize-ensurer.reset-size7.size5,.katex .katex-sizing.reset-size7.size5{font-size:.75em}.katex .fontsize-ensurer.reset-size7.size6,.katex .katex-sizing.reset-size7.size6{font-size:.8333333333em}.katex .fontsize-ensurer.reset-size7.size7,.katex .katex-sizing.reset-size7.size7{font-size:1em}.katex .fontsize-ensurer.reset-size7.size8,.katex .katex-sizing.reset-size7.size8{font-size:1.2em}.katex .fontsize-ensurer.reset-size7.size9,.katex .katex-sizing.reset-size7.size9{font-size:1.44em}.katex .fontsize-ensurer.reset-size7.size10,.katex .katex-sizing.reset-size7.size10{font-size:1.7283333333em}.katex .fontsize-ensurer.reset-size7.size11,.katex .katex-sizing.reset-size7.size11{font-size:2.0733333333em}.katex .fontsize-ensurer.reset-size8.size1,.katex .katex-sizing.reset-size8.size1{font-size:.3472222222em}.katex .fontsize-ensurer.reset-size8.size2,.katex .katex-sizing.reset-size8.size2{font-size:.4166666667em}.katex .fontsize-ensurer.reset-size8.size3,.katex .katex-sizing.reset-size8.size3{font-size:.4861111111em}.katex .fontsize-ensurer.reset-size8.size4,.katex .katex-sizing.reset-size8.size4{font-size:.5555555556em}.katex .fontsize-ensurer.reset-size8.size5,.katex .katex-sizing.reset-size8.size5{font-size:.625em}.katex .fontsize-ensurer.reset-size8.size6,.katex .katex-sizing.reset-size8.size6{font-size:.6944444444em}.katex .fontsize-ensurer.reset-size8.size7,.katex .katex-sizing.reset-size8.size7{font-size:.8333333333em}.katex .fontsize-ensurer.reset-size8.size8,.katex .katex-sizing.reset-size8.size8{font-size:1em}.katex .fontsize-ensurer.reset-size8.size9,.katex .katex-sizing.reset-size8.size9{font-size:1.2em}.katex .fontsize-ensurer.reset-size8.size10,.katex .katex-sizing.reset-size8.size10{font-size:1.4402777778em}.katex .fontsize-ensurer.reset-size8.size11,.katex .katex-sizing.reset-size8.size11{font-size:1.7277777778em}.katex .fontsize-ensurer.reset-size9.size1,.katex .katex-sizing.reset-size9.size1{font-size:.2893518519em}.katex .fontsize-ensurer.reset-size9.size2,.katex .katex-sizing.reset-size9.size2{font-size:.3472222222em}.katex .fontsize-ensurer.reset-size9.size3,.katex .katex-sizing.reset-size9.size3{font-size:.4050925926em}.katex .fontsize-ensurer.reset-size9.size4,.katex .katex-sizing.reset-size9.size4{font-size:.462962963em}.katex .fontsize-ensurer.reset-size9.size5,.katex .katex-sizing.reset-size9.size5{font-size:.5208333333em}.katex .fontsize-ensurer.reset-size9.size6,.katex .katex-sizing.reset-size9.size6{font-size:.5787037037em}.katex .fontsize-ensurer.reset-size9.size7,.katex .katex-sizing.reset-size9.size7{font-size:.6944444444em}.katex .fontsize-ensurer.reset-size9.size8,.katex .katex-sizing.reset-size9.size8{font-size:.8333333333em}.katex .fontsize-ensurer.reset-size9.size9,.katex .katex-sizing.reset-size9.size9{font-size:1em}.katex .fontsize-ensurer.reset-size9.size10,.katex .katex-sizing.reset-size9.size10{font-size:1.2002314815em}.katex .fontsize-ensurer.reset-size9.size11,.katex .katex-sizing.reset-size9.size11{font-size:1.4398148148em}.katex .fontsize-ensurer.reset-size10.size1,.katex .katex-sizing.reset-size10.size1{font-size:.2410800386em}.katex .fontsize-ensurer.reset-size10.size2,.katex .katex-sizing.reset-size10.size2{font-size:.2892960463em}.katex .fontsize-ensurer.reset-size10.size3,.katex .katex-sizing.reset-size10.size3{font-size:.337512054em}.katex .fontsize-ensurer.reset-size10.size4,.katex .katex-sizing.reset-size10.size4{font-size:.3857280617em}.katex .fontsize-ensurer.reset-size10.size5,.katex .katex-sizing.reset-size10.size5{font-size:.4339440694em}.katex .fontsize-ensurer.reset-size10.size6,.katex .katex-sizing.reset-size10.size6{font-size:.4821600771em}.katex .fontsize-ensurer.reset-size10.size7,.katex .katex-sizing.reset-size10.size7{font-size:.5785920926em}.katex .fontsize-ensurer.reset-size10.size8,.katex .katex-sizing.reset-size10.size8{font-size:.6943105111em}.katex .fontsize-ensurer.reset-size10.size9,.katex .katex-sizing.reset-size10.size9{font-size:.8331726133em}.katex .fontsize-ensurer.reset-size10.size10,.katex .katex-sizing.reset-size10.size10{font-size:1em}.katex .fontsize-ensurer.reset-size10.size11,.katex .katex-sizing.reset-size10.size11{font-size:1.1996142719em}.katex .fontsize-ensurer.reset-size11.size1,.katex .katex-sizing.reset-size11.size1{font-size:.2009646302em}.katex .fontsize-ensurer.reset-size11.size2,.katex .katex-sizing.reset-size11.size2{font-size:.2411575563em}.katex .fontsize-ensurer.reset-size11.size3,.katex .katex-sizing.reset-size11.size3{font-size:.2813504823em}.katex .fontsize-ensurer.reset-size11.size4,.katex .katex-sizing.reset-size11.size4{font-size:.3215434084em}.katex .fontsize-ensurer.reset-size11.size5,.katex .katex-sizing.reset-size11.size5{font-size:.3617363344em}.katex .fontsize-ensurer.reset-size11.size6,.katex .katex-sizing.reset-size11.size6{font-size:.4019292605em}.katex .fontsize-ensurer.reset-size11.size7,.katex .katex-sizing.reset-size11.size7{font-size:.4823151125em}.katex .fontsize-ensurer.reset-size11.size8,.katex .katex-sizing.reset-size11.size8{font-size:.578778135em}.katex .fontsize-ensurer.reset-size11.size9,.katex .katex-sizing.reset-size11.size9{font-size:.6945337621em}.katex .fontsize-ensurer.reset-size11.size10,.katex .katex-sizing.reset-size11.size10{font-size:.8336012862em}.katex .fontsize-ensurer.reset-size11.size11,.katex .katex-sizing.reset-size11.size11{font-size:1em}.katex .delimsizing.size1{font-family:KaTeX_Size1}.katex .delimsizing.size2{font-family:KaTeX_Size2}.katex .delimsizing.size3{font-family:KaTeX_Size3}.katex .delimsizing.size4{font-family:KaTeX_Size4}.katex .delimsizing.mult .delim-size1>span{font-family:KaTeX_Size1}.katex .delimsizing.mult .delim-size4>span{font-family:KaTeX_Size4}.katex .nulldelimiter{display:inline-block;width:.12em}.katex .delimcenter,.katex .op-symbol{position:relative}.katex .op-symbol.small-op{font-family:KaTeX_Size1}.katex .op-symbol.large-op{font-family:KaTeX_Size2}.katex .katex-accent>.vlist-t,.katex .op-limits>.vlist-t{text-align:center}.katex .katex-accent .accent-body{position:relative}.katex .katex-accent .accent-body:not(.accent-full){width:0}.katex .katex-overlay{display:block}.katex .mtable .vertical-separator{display:inline-block;min-width:1px}.katex .mtable .arraycolsep{display:inline-block}.katex .mtable .col-align-c>.vlist-t{text-align:center}.katex .mtable .col-align-l>.vlist-t{text-align:left}.katex .mtable .col-align-r>.vlist-t{text-align:right}.katex .svg-align{text-align:left}.katex svg{fill:currentColor;stroke:currentColor;display:block;height:inherit;position:absolute;width:100%}.katex svg path{stroke:none}.katex svg{fill-rule:nonzero;fill-opacity:1;stroke-width:1;stroke-linecap:butt;stroke-linejoin:miter;stroke-miterlimit:4;stroke-dasharray:none;stroke-dashoffset:0;stroke-opacity:1}.katex img{border-style:none;max-height:none;max-width:none;min-height:0;min-width:0}.katex .katex-stretchy{display:block;overflow:hidden;position:relative;width:100%}.katex .katex-stretchy:after,.katex .katex-stretchy:before{content:""}.katex .hide-tail{overflow:hidden;position:relative;width:100%}.katex .halfarrow-left{left:0;overflow:hidden;position:absolute;width:50.2%}.katex .halfarrow-right{overflow:hidden;position:absolute;right:0;width:50.2%}.katex .brace-left{left:0;overflow:hidden;position:absolute;width:25.1%}.katex .brace-center{left:25%;overflow:hidden;position:absolute;width:50%}.katex .brace-right{overflow:hidden;position:absolute;right:0;width:25.1%}.katex .x-arrow-pad{padding:0 .5em}.katex .cd-arrow-pad{padding:0 .55556em 0 .27778em}.katex .mover,.katex .munder,.katex .x-arrow{text-align:center}.katex .boxpad{padding:0 .3em}.katex .fbox,.katex .fcolorbox{border:.04em solid;box-sizing:border-box}.katex .cancel-pad{padding:0 .2em}.katex .cancel-lap{margin-left:-.2em;margin-right:-.2em}.katex .katex-sout{border-bottom-style:solid;border-bottom-width:.08em}.katex .angl{border-right:.049em solid;border-top:.049em solid;box-sizing:border-box;margin-right:.03889em}.katex .anglpad{padding:0 .03889em}.katex .reflectbox{display:inline-block;transform:scaleX(-1)}.katex .eqn-num:before{content:"(" counter(katexEqnNo) ")";counter-increment:katexEqnNo}.katex .mml-eqn-num:before{content:"(" counter(mmlEqnNo) ")";counter-increment:mmlEqnNo}.katex .mtr-glue{width:50%}.katex .cd-vert-arrow{display:inline-block;position:relative}.katex .cd-label-left{display:inline-block;position:absolute;right:calc(50% + .3em);text-align:left}.katex .cd-label-right{display:inline-block;left:calc(50% + .3em);position:absolute;text-align:right}.katex-display{display:block;margin:1em 0;text-align:center}.katex-display>.katex{display:block;text-align:center;white-space:nowrap}.katex-display>.katex>.katex-html{display:block;position:relative}.katex-display>.katex>.katex-html>.katex-tag{position:absolute;right:0}.katex-display.leqno>.katex>.katex-html>.katex-tag{left:0;right:auto}.katex-display.fleqn>.katex{padding-left:2em;text-align:left}body{counter-reset:katexEqnNo mmlEqnNo}
</style>
<script>(function() {
function __mumInit() {
// Restore external links that htmlpreview.github.io's loader rewrote to
// in-page anchors (it treats any href containing '#' as a local anchor).
// Runs now (htmlpreview re-executes this script after its rewrite), again
// after delays, and on click as a last line of defense.
function restoreOrigHrefs() {
document.querySelectorAll('a[data-orig-href]').forEach(function(a) {
var orig = a.getAttribute('data-orig-href');
if (orig && a.getAttribute('href') !== orig) a.setAttribute('href', orig);
});
}
restoreOrigHrefs();
setTimeout(restoreOrigHrefs, 100);
setTimeout(restoreOrigHrefs, 1000);
document.addEventListener('click', function(e) {
var t = e.target && e.target.closest ? e.target.closest('a[data-orig-href]') : null;
if (t) { var o = t.getAttribute('data-orig-href'); if (o && t.getAttribute('href') !== o) t.setAttribute('href', o); }
}, true);
var vt = document.getElementById('viewing-time');
if (vt) {
function pad(n) { return n < 10 ? '0' + n : n; }
function updateTime() {
var d = new Date();
vt.textContent = d.getUTCFullYear() + '-' + pad(d.getUTCMonth()+1) + '-' + pad(d.getUTCDate()) + ' ' + pad(d.getUTCHours()) + ':' + pad(d.getUTCMinutes()) + ' UTC';
}
updateTime();
setInterval(updateTime, 60000);
}
document.querySelectorAll('a[data-hover-title]').forEach(function(a) {
var card = document.createElement('span');
card.className = 'hover-card';
var title = a.getAttribute('data-hover-title') || '';
var status = a.getAttribute('data-hover-status') || '';
var owner = a.getAttribute('data-hover-owner') || '';
var desc = a.getAttribute('data-hover-desc') || '';
var statusCls = 'hc-status hc-status-' + status.toLowerCase().replace(/[^a-z]/g, '');
var html = '<div class="hc-title">' + title + '</div>';
var meta = [];
if (status) meta.push('<span class="' + statusCls + '">' + status + '</span>');
if (owner) meta.push(owner);
if (meta.length) html += '<div class="hc-meta">' + meta.join(' ') + '</div>';
if (desc) html += '<div class="hc-desc">' + desc + '</div>';
card.innerHTML = html;
a.appendChild(card);
});
document.querySelectorAll('a.cite').forEach(function(a) {
var href = a.getAttribute('href') || '';
if (!href.startsWith('#ref-')) return;
var li = document.getElementById(href.slice(1));
if (!li) return;
var body = li.querySelector('.ref-body');
if (!body) return;
var refUrl = body.querySelector('.ref-url a');
var refHref = refUrl ? refUrl.getAttribute('href') : '';
var hoverSource = refHref ? document.querySelector('a[data-hover-title][href="' + refHref + '"]') : null;
if (hoverSource) {
var card = document.createElement('span');
card.className = 'hover-card';
var t = hoverSource.getAttribute('data-hover-title') || '';
var s = hoverSource.getAttribute('data-hover-status') || '';
var o = hoverSource.getAttribute('data-hover-owner') || '';
var d = hoverSource.getAttribute('data-hover-desc') || '';
var sc = 'hc-status hc-status-' + s.toLowerCase().replace(/[^a-z]/g, '');
var h = '<div class="hc-title">' + t + '</div>';
var m = [];
if (s) m.push('<span class="' + sc + '">' + s + '</span>');
if (o) m.push(o);
if (m.length) h += '<div class="hc-meta">' + m.join(' ') + '</div>';
if (d) h += '<div class="hc-desc">' + d + '</div>';
card.innerHTML = h;
a.appendChild(card);
} else {
var title = body.querySelector('.ref-title');
if (!title) return;
var card = document.createElement('span');
card.className = 'cite-card';
card.textContent = title.textContent;
a.appendChild(card);
}
});
}
if (document.readyState !== 'loading') { __mumInit(); }
else { document.addEventListener('DOMContentLoaded', __mumInit); }
})();</script>
</head>
<body>
<div class="page">
<header class="paper-header">
<h1>Covering the competency map and the task curriculum: where Marin&#x27;s RL environments come from</h1>
<div class="paper-meta">Posed by <span class="author">hammer</span>, answered by <a class="mumwelt-link" href="https://github.com/marin-community/mumwelt">mumwelt</a> &middot; <span class="date-label" title="Generated 2026-09-22 06:03 UTC">Published 2026-09-22</span></div>
</header>
<div class="prompt-box">
<div class="prompt-label">?</div>
<div class="prompt-text">How is Marin generating RL environments to cover its competency taxonomy and task curriculum? What does it mean to cover the taxonomy and curriculum, and specifically how is GLM-5.3 being used to generate an environment for each element of them?</div>
</div>
<div class="content">
<p><em>State of play as of September 22, 2026, from Marin's GitHub,
Discord, weekly summaries, and the linked design documents.</em></p>
<h2 id="short-answer">Short answer</h2>
<p>Marin has two different maps, built by two different people with two
different models, and neither one yet has an environment generated per
element.</p>
<ol type="1">
<li><strong>The competency coverage map</strong> (Benjamin Feuer,
September 3) takes nine public taxonomies of knowledge and work, boils
them down to 151 trainable "competencies," and checks which ones
TaskTrove's 1.76 million tasks already touch. Twenty-three are covered.
One hundred twenty-six have no task at all. Its purpose is to point at
gaps (<a
href="https://github.com/marin-community/marin/issues/8879">#8879</a>,
<a
href="https://storage.googleapis.com/marin-public/benjaminfeuer/tasktrove-competency-coverage/2026.09.03/index.html">report</a>).</li>
<li><strong>The cross-domain task curriculum</strong> (Russell Power,
September 17 to 20) is a separate, finer catalog: 45 subjects, 1,999
trainable "capabilities," and a reviewed graph of 1,812 prerequisite
edges. It was generated and reviewed by OpenAI's
<code>gpt-5.6-sol</code> running under Codex, with a cheaper "Luna" tier
for blind checks (<a
href="https://github.com/marin-community/marin/pull/9232">#9232</a>, <a
href="https://github.com/marin-community/marin/pull/9294">#9294</a>, <a
href="https://github.com/marin-community/marin/blob/main/experiments/post_training/task_curriculum/workflow.md">workflow</a>).
GLM-5.3 played no part in building it.</li>
</ol>
<p>GLM-5.3 is Marin's workhorse open-weight model for post-training
data, and it does four jobs in this area. The one closest to "an
environment per element" is Mark Muchane's pipeline for the 126 gap
competencies, where a GLM-5.3 agent extracts release histories from
practitioner repositories that will later be turned into tasks. That
pipeline was still at the "extract and propose" stage in mid-September,
and no environments from it have been published or tracked in an issue
yet. The other three jobs are routing existing TaskTrove tasks,
producing teacher rollouts for supervised fine-tuning, and producing
baselines for a biology track.</p>
<p>The rest of this note explains what "covered" means in each map, what
an RL environment is in Marin's stack today, what GLM-5.3 actually does,
and what has been shown to work.</p>
<h2 id="two-maps-two-meanings-of-covered">Two maps, two meanings of
"covered"</h2>
<h3 id="the-coverage-map-a-gap-finder">The coverage map: a gap
finder</h3>
<p>Feuer's audit normalized 85,354 nodes from OpenAlex, arXiv, ANZSRC,
OECD Fields of R&amp;D, MSC 2020, MeSH, NAICS, and the O*NET occupation
and work-activity hierarchies into one parent-child schema, then picked
5,497 mid-level units to avoid comparing 65,000 fine-grained MeSH leaves
with much coarser taxonomies. He hand-wrote 35 macro and 160 micro
competencies, each of which "had to describe recurring, observable
computer work with a plausible verification method," and assigned every
taxonomy unit to one to three of them with a confidence score. Nine
received no unit, leaving 151 (<a
href="https://storage.googleapis.com/marin-public/benjaminfeuer/tasktrove-competency-coverage/2026.09.03/index.html">report</a>).</p>
<p>Each of TaskTrove's 96 sources was then mapped to the same
vocabulary. "Covered" is a screening threshold: zero tasks is a gap,
under 250 is very sparse, 250 to 999 is partial, and 1,000 or more is
covered. The issue is explicit that this "prioritizes candidate bounty
areas; it does not establish training sufficiency," and that a source's
whole row count is credited to every competency it maps to, so the
counts are corpus volume rather than a partition of unique tasks (<a
href="https://github.com/marin-community/marin/issues/8879">#8879</a>).
The named gaps are healthcare workflows, data analysis, office
documents, scientific computing, research workflows, engineering design,
finance operations, and business administration.</p>
<p>Feuer announced it on Discord with an open call: <a
href="https://discord.com/channels/1354881461060243556/1368297424086499359/1545147565496999979">"Pick
a knowledge domain where we don't have coverage, submit a PR with a new
dataset that introduces an RL dataset!"</a>. Asked whether any gaps had
priority, he answered that <a
href="https://discord.com/channels/1354881461060243556/1545147773391732758/1545239773419929741">"anything
is welcome, but we are always particularly interested in science
data!"</a>. No formal bounty list followed. The issue has no comments.
Contributors self-selected, as described below.</p>
<h3 id="the-task-curriculum-a-capability-catalog">The task curriculum: a
capability catalog</h3>
<p>Power's curriculum starts from a different inventory, 45 subject
roots with 356 "guideposts," and uses ISCED-F, CIP, Frascati, and
occupational taxonomies only as coverage checks. Its README says its job
is to turn "a broad subject inventory into trainable curricula and maps
tasks onto reviewed sections," and to keep curriculum design separate
from task correctness: "TaskCompendium owns model-visible task semantics
and private verifier contracts; a curriculum describes observable
capabilities, boundaries, and examples" (<a
href="https://github.com/marin-community/marin/blob/main/experiments/post_training/task_curriculum/README.md">README</a>).
The README lists Feuer's coverage audit as an input, so the two maps are
related in lineage but not in structure.</p>
<p>The 45 subjects run from D01 Mathematics &amp; Statistics and D02
Computer &amp; Information Sciences through D17 Medicine &amp; Clinical
Care, D27 Finance, Accounting &amp; Audit, D28 Law, Regulation &amp;
Criminology, to D45 Skilled Trades, Fabrication, Maintenance &amp;
Repair. D02 is the largest with 95 capabilities in 13 groups (<a
href="https://public.applets.marina.oa.dev/a/67f69132-2ef4-4c9e-b8b5-77cabd126442/v/9/">public
viewer, revision 9</a>). A capability is the only node kind that may
receive task assignments. Each has an outcome statement, includes and
excludes lists, and sampling facets. For example,
<code>d02.pl.parsing</code> has the outcome "Transform a lexical and
grammatical specification into a parser that produces the required
syntax structure and rejects malformed programs at the correct
boundary," and excludes "repository-level defect localization" (<a
href="https://public.applets.marina.oa.dev/a/67f69132-2ef4-4c9e-b8b5-77cabd126442/v/9/api/catalog">catalog
JSON</a>).</p>
<p>How it was built matters for the question you asked. The workflow
document states that "the reference role model is
<code>gpt-5.6-sol</code> with high reasoning effort" (<a
href="https://github.com/marin-community/marin/blob/main/experiments/post_training/task_curriculum/workflow.md">workflow.md</a>);
"Sol" is OpenAI's GPT-5.6 tier run through Codex (<a
data-orig-href="https://github.com/marin-community/marin/issues/7323#issuecomment-5006373740" href="https://github.com/marin-community/marin/issues/7323#issuecomment-5006373740">#7323</a>).
Each subject got one Sol/high proposer and one independent Sol/high
reviewer working from a compact packet; agents never saw the full
catalog. The original build cost 16.28 million tokens, a repair wave
8.92 million, and the progression pass 4.85 million (<a
href="https://github.com/marin-community/marin/pull/9294">#9294</a>).
Repairs were accepted only when a fresh anonymous reviewer preferred the
repaired graph and scored it at least 70; the mean holistic score rose
from 78.36 to 92.81, with 24 subjects marked <code>pilot_ready</code>
and 21 left "explicitly provisional" (<a
href="https://github.com/marin-community/marin/pull/9232">#9232</a>).
The "capability scores" tab in the viewer is this audit table, not model
performance. The 1,812 prerequisite edges implement one rule, that
"mastery of A should materially improve the chance of some success on a
recurring family of entry-level B tasks," and the README calls them
"structural learning hypotheses, not causal transfer measurements" (<a
href="https://github.com/marin-community/marin/blob/main/experiments/post_training/task_curriculum/README.md">README</a>).</p>
<p>Both PRs were authored and merged by Power with only bot reviewers on
the record (<a
href="https://github.com/marin-community/marin/pull/9232">#9232
reviews</a>, <a
href="https://github.com/marin-community/marin/pull/9294">#9294</a>). No
RL run consumes the catalog yet. The only in-repo consumer is a mapping
tool that embeds a task's semantic key and places it on a capability,
with "roughly 70% reasonable placement" deemed adequate and mapping
"excluded from the curriculum promotion gate" (<a
href="https://github.com/marin-community/marin/blob/cf3f4ffd1b8821a2216eac2082409306c2a29cb2/experiments/post_training/task_curriculum/task_mapping/cli.py">task_mapping</a>).
A routing audit placed all 24 sampled TaskTrove shell, repository, and
bug-repair tasks in D02 (<a
href="https://github.com/marin-community/marin/pull/9232">#9232</a>).
Power's stated goal for all this tagging is to have "enough latent
information + information from a task to be able to express a desired RL
data mixture" (<a
data-orig-href="https://github.com/marin-community/marin/pull/9187#issuecomment-5690740119" href="https://github.com/marin-community/marin/pull/9187#issuecomment-5690740119">review
on #9187</a>), and he expects TaskTrove's grab-bag sources will need to
be "blown up to align with the curriculum categories" (<a
data-orig-href="https://github.com/marin-community/marin/pull/9264#issuecomment-5743127086" href="https://github.com/marin-community/marin/pull/9264#issuecomment-5743127086">#9264
comment</a>).</p>
<p>So "covering the curriculum" today means: every capability has a
definition and a place in a prerequisite graph, and there is a tool to
route tasks onto it. It does not yet mean every capability has tasks,
and it does not mean a model has been trained against it.</p>
<h2 id="what-an-rl-environment-is-in-marin-right-now">What an RL
environment is in Marin right now</h2>
<p>A trainable task in Marin is a Harbor task directory: an
<code>instruction.md</code>, a shared Dockerfile with no task-specific
data, a hidden <code>tests/</code> folder holding a typed grader
contract in <code>verifier.toml</code>, and optionally a reference
solution. Power's TaskTrove conversion pipeline converts upstream rows
into this form and rejects anything that "need[s] answer-recovery
heuristics, repository reconstruction, replacement tests, or a new
subjective contract." Release 2026.09.10.9 kept 1,449,686 of 1,739,326
rows across 43 sources, 12 verifier modes, and 39 Docker environments,
with every rejection logged (<a
href="https://github.com/marin-community/marin/pull/9061">#9061</a>, <a
href="https://github.com/marin-community/marin/blob/cf3f4ffd1b8821a2216eac2082409306c2a29cb2/experiments/post_training/tasktrove/verify.py">verify.py</a>).
A missing verdict is scored as "never a pass" (<a
data-orig-href="https://github.com/marin-community/marin/pull/9061#issuecomment-5627275993" href="https://github.com/marin-community/marin/pull/9061#issuecomment-5627275993">#9061
comment</a>).</p>
<p>The environment contract issue makes identity and reuse explicit: "An
environment identity includes the Dockerfile, pinned verifier, and
sandbox runtime interface. Equivalent environments reuse snapshots
across runs." Tasks must declare capabilities they need, such as network
egress or judge access, and validation should require that an empty
workspace scores 0 and the golden solution scores 1 (<a
href="https://github.com/marin-community/marin/issues/9122">#9122</a>,
<a
href="https://github.com/marin-community/marin/issues/9083">#9083</a>).
SkyRL experiments select rows "by exact source, tags, modes, limit, and
seed" from one packed Parquet (<a
href="https://github.com/marin-community/marin/pull/9125">#9125</a>).</p>
<p>David Hall's TaskCompendium is the successor format. It <a
href="https://github.com/marin-community/marin/pull/9187">"separates a
task's semantics and private verifier from its submission format and
harness,"</a> so the same semantic task can be lowered to a chat prompt,
a Harbor sandbox, or a NeMo tool-call check. The spike holds 63
specifications and 164 Harbor lowerings and carries its own tag ontology
of 17 subject roots and a competency namespace, which Hall calls
"definitely not baked... a POC" (<a
href="https://github.com/marin-community/marin/pull/9187">#9187</a>, <a
data-orig-href="https://github.com/marin-community/marin/pull/9187#issuecomment-5700804140" href="https://github.com/marin-community/marin/pull/9187#issuecomment-5700804140">reply</a>,
<a href="https://github.com/marin-community/marin/pull/9214">#9214</a>).
In the September 15 standup Hall described TaskCompendium as an
amalgamation of a "Coverage-Directed Agentic Environment Generation"
design and Power's "Unified Task Format" design (<a
href="https://github.com/marin-community/marin/issues/9171">#9171</a>).
Those two documents are access-restricted, so this note cannot quote
them.</p>
<h2 id="how-glm-53-is-actually-used">How GLM-5.3 is actually used</h2>
<p>GLM-5.3 is zai-org's open-weight model. Mark Muchane runs a
self-hosted inference pool for it and wrote in the September 15 standup
that "We have capacity for hundreds of Opus 4.8-strength models running
at very fast throughput (generally ~200tok/s+)" (<a
href="https://github.com/marin-community/marin/issues/9171">#9171</a>).
Access is by request; he offered to get a collaborator "added to the
cluster where the GLM endpoint lives" (<a
href="https://discord.com/channels/1354881461060243556/1550321176381886535/1550614395900665887">Discord</a>).
Marin's preference for GLM-family teachers dates to OpenThoughts-Agent
ablations in which "GLM-4.6 produces 2x stronger Qwen3-8B students than
GPT-5, despite being 10 points behind on TerminalBench" (<a
href="https://discord.com/channels/1354881461060243556/1476636771675541525/1476637266335109338">Kevin
Xiang Li</a>), and to a judge study that found GLM-5.1 within about
0.009 median Spearman of GPT-5.1 at roughly 16 times lower cost (<a
href="https://github.com/marin-community/marin/issues/4790">#4790</a>).</p>
<h3 id="job-1-building-environments-for-the-126-gap-competencies">Job 1:
building environments for the 126 gap competencies</h3>
<p>This is the work your question is really about. Muchane announced on
September 9 that he was "working on synthetic task generation with GLM
5.3 for areas we're missing" and shared a design doc, "Building Agentic
Environments for Missing TaskTrove Competencies" (<a
href="https://discord.com/channels/1354881461060243556/1368297424086499359/1547373328728334377">Discord</a>,
<a
href="https://docs.google.com/document/d/1ED4ThLOieKT_XIPNSUiL_7Sghnn6bhSwrTfeGNhYHqY/edit">doc</a>).
The doc starts from Feuer's 126 gaps and proposes building agentic
environments for the ones that <a
href="https://docs.google.com/document/d/1ED4ThLOieKT_XIPNSUiL_7Sghnn6bhSwrTfeGNhYHqY/edit">"make
sense as being defined as a coding task"</a> and <a
href="https://docs.google.com/document/d/1ED4ThLOieKT_XIPNSUiL_7Sghnn6bhSwrTfeGNhYHqY/edit">"have
open-source repositories that are used by practitioners in the
competencies."</a> Success is defined as environments that cover the
competencies that make sense, that are <a
href="https://docs.google.com/document/d/1ED4ThLOieKT_XIPNSUiL_7Sghnn6bhSwrTfeGNhYHqY/edit">"not
broken and do not encourage reward hacking,"</a> and that make a model
<a
href="https://docs.google.com/document/d/1ED4ThLOieKT_XIPNSUiL_7Sghnn6bhSwrTfeGNhYHqY/edit">"better
on some metric we care about (ideally better faster than a sample of the
closest matching existing TaskTrove environments)."</a></p>
<p>The pipeline, per competency:</p>
<ul>
<li><strong>Stage 0, repository discovery.</strong> Codex with web
search finds 10 to 50 practitioner-facing repositories per competency,
preferring "substantial self-contained functionality" over glue code,
and must open each repository page before including it (<a
href="https://pastebin.com/dZHVXugJ">prompt</a>). The example for
"Accessibility and Responsive Compliance" returned 19 repositories such
as axe-core, NVDA, Lighthouse, and Playwright (<a
href="https://pastebin.com/PPKxXQEL">list</a>). Muchane notes in a doc
comment that Kimi with a web-search tool <a
href="https://docs.google.com/document/d/1ED4ThLOieKT_XIPNSUiL_7Sghnn6bhSwrTfeGNhYHqY/edit">"did
just as well as codex."</a></li>
<li><strong>Stage 1, release-history extraction.</strong> This is where
GLM-5.3 does the work. For each repository, a GLM-5.3 agent running in
the open-source Oh My Pi coding agent (<a
href="https://github.com/can1357/oh-my-pi">omp</a>) must work out how
the project publishes releases, write a re-runnable
<code>recipe.sh</code>, and emit one JSON record per release with the
original release text preserved verbatim. The prompt lists a dozen
silent-failure traps, such as GitHub's prerelease flag being wrong in
both directions and changelogs that are years stale, and forbids
synthesizing notes from commit logs because "that produces a different
kind of artifact and would silently contaminate the corpus" (<a
href="https://pastebin.com/CQqXR7b1">prompt</a>). The example extraction
for Siteimprove/alfa recovered 213 releases (<a
href="https://pastebin.com/h0mgeybA">example</a>).</li>
<li><strong>Stage 1a, select releases.</strong> Roughly 50 per
repository, favoring feature releases spread over time. Undecided.</li>
<li><strong>Stage 2, graded task proposals.</strong> Each
release-repository pair becomes proposals for tasks, graded on how well
the repository authors already defined the work, how self-contained it
is, how relevant it is to the competency rather than to the library, and
how well it fits a lightweight sandbox. Bug fixes become <a
href="https://docs.google.com/document/d/1ED4ThLOieKT_XIPNSUiL_7Sghnn6bhSwrTfeGNhYHqY/edit">"find
the bug and fix it"</a> tasks; documentation updates become retrieval
tasks.</li>
<li><strong>Stages 3 and 4, generate and evaluate tasks.</strong> Marked
<a
href="https://docs.google.com/document/d/1ED4ThLOieKT_XIPNSUiL_7Sghnn6bhSwrTfeGNhYHqY/edit">"Haven't
worked this out"</a> in the doc, with a plan to evaluate by running
rollouts to catch reward hacking and measure difficulty.</li>
</ul>
<p>Inline reviewer comments on the doc changed the plan in three ways.
The original aim to run environments in Power's in-process shell
simulator (<a href="https://github.com/rjpower/shellsim">shellsim</a>)
drew pushback (<a
href="https://docs.google.com/document/d/1ED4ThLOieKT_XIPNSUiL_7Sghnn6bhSwrTfeGNhYHqY/edit">"Why
would we want to systematically distort our RL task distribution to
match a transient piece of infra?"</a>) and a reminder that Marin has <a
href="https://docs.google.com/document/d/1ED4ThLOieKT_XIPNSUiL_7Sghnn6bhSwrTfeGNhYHqY/edit">"literally
10,000s of free CPUs on Daytona"</a>; Muchane replied that <a
href="https://docs.google.com/document/d/1ED4ThLOieKT_XIPNSUiL_7Sghnn6bhSwrTfeGNhYHqY/edit">"my
current work is no longer anchored on ShellSim since we have the daytona
CPUs."</a> Reviewers warned that <a
href="https://docs.google.com/document/d/1ED4ThLOieKT_XIPNSUiL_7Sghnn6bhSwrTfeGNhYHqY/edit">"agentic
!= coding"</a> and listed law, medicine, spreadsheets, and visualization
as agentic areas the repository-driven method misses. And Muchane
reported that <a
href="https://docs.google.com/document/d/1ED4ThLOieKT_XIPNSUiL_7Sghnn6bhSwrTfeGNhYHqY/edit">"GLM
is not super amazing at evaluating tasks from just a proposal,"</a> so
the grading axes will move after environment construction or first
rollouts (<a
href="https://docs.google.com/document/d/1ED4ThLOieKT_XIPNSUiL_7Sghnn6bhSwrTfeGNhYHqY/edit">doc
comments</a>). One reviewer asked that human experts review <a
href="https://docs.google.com/document/d/1ED4ThLOieKT_XIPNSUiL_7Sghnn6bhSwrTfeGNhYHqY/edit">"our
synthetic expert-specialized RL data,"</a> at least for the hero
run.</p>
<p>Status: the September 15 standup lists <a
href="https://github.com/marin-community/marin/issues/9171">"Building a
synthetic RL env creation pipeline"</a> as last week's work and <a
href="https://github.com/marin-community/marin/issues/9171">"More RL env
data pipeline work"</a> as this week's. The doc estimated a week to get
the pipeline operating and said generation <a
href="https://docs.google.com/document/d/1ED4ThLOieKT_XIPNSUiL_7Sghnn6bhSwrTfeGNhYHqY/edit">"can
continue for as long as we have B200s (currently end of month)."</a> As
of the September 22 mirror there is no GitHub issue or PR tracking this
pipeline, no code for it in the Marin repository, and no published
environment count. Feuer has separately set a rule for community
contributions that bears on it: "let's avoid non reproducible (LLM
directly intermediated) synthetic task generation for now please? Better
to have the agents write a task template and embed that in a Python
script" (<a
href="https://discord.com/channels/1354881461060243556/1547218837404131359/1547258378093596734">Discord</a>).
That was addressed to a Kaggle-style proposal, not to Muchane, and
Muchane's Stage 1 prompt already demands a re-runnable script; whether
Stage 3 will meet the same bar is not yet visible.</p>
<h3 id="job-2-routing-tasktroves-multiple-choice-tasks">Job 2: routing
TaskTrove's multiple-choice tasks</h3>
<p>The one merged pipeline step that uses GLM-5.3 is Power's MCQA
routing pass. TaskTrove's cleaned release held 611,698 multiple-choice
tasks, most unsuited to RL. A reusable step <a
href="https://github.com/marin-community/marin/pull/9264">"submits
recoverable GLM-5.3 batches, validates a forced tool response
locally,"</a> and produced 603,290 decisions: 23,860 to RL, 458,562 to
SFT, and 120,868 garbage. Unmapped rows fail closed. The 2026.09.18.3
release keeps 861,848 tasks in the main corpus and adds separate
<code>rl/</code> and <code>sft/</code> slices; the default training
mixture did not change (<a
href="https://github.com/marin-community/marin/pull/9264">#9264</a>).
This is GLM-5.3 as a classifier over existing tasks, not a
generator.</p>
<h3 id="job-3-teacher-rollouts-and-sft-data">Job 3: teacher rollouts and
SFT data</h3>
<p>Alex Dimakis's group, building the coding expert, found that the
Snowball RL starting checkpoint <a
href="https://discord.com/channels/1354881461060243556/1550321176381886535/1550609642114130021">"solves
~5 tasks / 87"</a> on Terminal-Bench 2 and argued for <a
href="https://discord.com/channels/1354881461060243556/1550321176381886535/1550609642114130021">"a
rigorous SFT step to lift this performance to at least ~20-30%"</a>
before RL, using GLM-5.3 as the stronger teacher on the 37,000-task
Recursive-Task-Synthesis set (<a
href="https://discord.com/channels/1354881461060243556/1550321176381886535/1550609642114130021">Harshit
Varma</a>). Muchane ran a 200-task pilot on September 19 (<a
href="https://discord.com/channels/1354881461060243556/1550321176381886535/1550684333109411921">Discord</a>)
and then the full set at eight trials per task, publishing <a
href="https://huggingface.co/datasets/open-athena/recursive-task-synthesis-glm-5.3-rollouts">open-athena/recursive-task-synthesis-glm-5.3-rollouts</a>
(<a
href="https://discord.com/channels/1354881461060243556/1550321176381886535/1551004503795437689">Discord</a>).
Luke Lee's analysis two days later is sobering: 14,410 rollouts of at
least five turns yielded 3,012 successes; <a
data-orig-href="https://github.com/marin-community/marin/issues/9225#issuecomment-5753734747" href="https://github.com/marin-community/marin/issues/9225#issuecomment-5753734747">"The
teacher (GLM-5.3) itself passes 0.09 of the RST tasks at one attempt
under this harness"</a>; and SFT on RST alone <a
data-orig-href="https://github.com/marin-community/marin/issues/9225#issuecomment-5753734747" href="https://github.com/marin-community/marin/issues/9225#issuecomment-5753734747">"is
not a terminal-skill dataset at this dose, it is a quitting
dataset,"</a> since 86 percent of rollouts end with the agent declaring
completion at a median of seven turns and 75 percent of those are
failures. Keeping only successes removes the harm but adds little (<a
data-orig-href="https://github.com/marin-community/marin/issues/9225#issuecomment-5753734747" href="https://github.com/marin-community/marin/issues/9225#issuecomment-5753734747">#9225
follow-up</a>).</p>
<p>Will Held also used GLM-5.3 to make two small SFT sets aimed at
Snowball's weak harness behavior: 10,000 context compactions from
AgentTrove traces and 9,800 WildChat prompts with format constraints and
verified completions (<a
href="https://discord.com/channels/1354881461060243556/1368297424086499359/1548432746655785040">compactions</a>,
<a
href="https://discord.com/channels/1354881461060243556/1368297424086499359/1548913355182448661">format</a>,
<a
href="https://github.com/marin-community/marin/pull/9161">#9161</a>).</p>
<h3 id="job-4-a-science-track">Job 4: a science track</h3>
<p>Gonzalo Benegas proposed synthetic Harbor tasks for computational
biology, "containing instructions, input data, an environment, a
reference solution, and a verifier," with GLM-5.3 named as a candidate
teacher. The generation recipe and pilot size are unselected (<a
href="https://github.com/marin-community/marin/issues/9257">#9257</a>).
His baselines put GLM-5.3 at 81.79 percent on MMLU biology and 101 of
187 graded on BixBench CLI (<a
href="https://github.com/marin-community/marin/issues/9258">#9258</a>).</p>
<h2 id="other-routes-to-filling-the-gaps">Other routes to filling the
gaps</h2>
<p>Several efforts target gap competencies without GLM-5.3 and without a
per-element pipeline:</p>
<ul>
<li><strong>Taxonomy-driven web search</strong> is one of four primary
TaskTrove pipelines the OpenThoughts-Next data breakout chose on
September 10: "start from a taxonomy (e.g. capability/domain), then
search for relevant documentation, repositories, etc. to construct tasks
and environments." The other three are agent skill files, recorded
terminal sessions, and git repositories via SETA (<a
href="https://discord.com/channels/1354881461060243556/1484315476325826660/1547671961851662346">Franziska
Weindel</a>). No PR has landed for any of the four.</li>
<li><strong>Data analysis and modeling.</strong> A contributor proposed
Kaggle-style tasks from OpenML, verifiable against held-out labels, and
opened a draft with samples (<a
href="https://discord.com/channels/1354881461060243556/1368297424086499359/1547218837404131359">proposal</a>,
<a
href="https://github.com/open-thoughts/OpenThoughts-Agent/pull/146">OpenThoughts-Agent
#146</a>).</li>
<li><strong>SQL engineering.</strong> A text-to-SQL dataset was routed
to MarinSkyRL as a native environment rather than TaskTrove (<a
href="https://discord.com/channels/1354881461060243556/1368297424086499359/1547004047100223608">proposal</a>,
<a
href="https://github.com/marin-community/MarinSkyRL/pull/539">MarinSkyRL
#539</a>).</li>
<li><strong>Executable repositories.</strong> Lena Lincke is auditing
SETA's roughly 5,000 tasks; Feuer's response was to ask whether the data
is "uniquely valuable" before adding it (<a
href="https://discord.com/channels/1354881461060243556/1547218837404131359/1549397917339615383">Discord</a>).</li>
</ul>
<p>None of these addresses healthcare workflows, office documents, or
finance operations, three of the eight named gaps.</p>
<h2 id="what-has-been-shown-to-work">What has been shown to work</h2>
<p>The honest summary is that environment breadth has not yet been shown
to move a headline metric, and the measured wins are in-distribution.
Twenty-four GRPO updates on R2E-Gym repo-fix tasks took Snowball from 12
to 23 on SWE-bench Verified random-100, while Terminal-Bench 2 stayed
flat; Lee's reading is that "The RL gain is in-distribution" (<a
href="https://github.com/marin-community/marin/issues/9225">#9225</a>).
Supervised fine-tuning on OpenThoughts-Agent traces reached the same
SWE-bench level with no RL at all (<a
data-orig-href="https://github.com/marin-community/marin/issues/9225#issuecomment-5753734747" href="https://github.com/marin-community/marin/issues/9225#issuecomment-5753734747">#9225
follow-up</a>). Varma's RL runs on broader synthetic pools showed
on-task gains but Terminal-Bench 2 changes within noise, and "88.5% of
RTS tasks are never solved in 8 tries" (<a
href="https://discord.com/channels/1354881461060243556/1551721992040882276/1551722141622206616">Discord</a>,
<a
href="https://discord.com/channels/1354881461060243556/1551721992040882276/1551722069488836719">Discord</a>).</p>
<p>Task quality is the other constraint. Hall's audit of 40 TaskTrove
tasks across 20 sources found "14/40 (35%) broken or hackable," 40
percent functional but low value, and 25 percent accepted (<a
href="https://discord.com/channels/1354881461060243556/1368297424086499359/1545694628756586587">Discord</a>).
That audit drove four TaskTrove point releases in one week and Power's
fail-closed conversion pipeline. Any new environment from the gap
pipeline will be judged by the same empty-workspace and golden-solution
gates.</p>
<p>The schedule these efforts serve: Feuer's roadmap puts the first
official Snowball post-train at "circa October 6" (<a
href="https://discord.com/channels/1354881461060243556/1354881461060243561/1549793034877673654">Discord</a>),
with expert contributions due "within the next 10-12 days," all data
public, and no traces from closed frontier models (<a
href="https://discord.com/channels/1354881461060243556/1374989195109466122/1549795035636047932">Discord</a>).
His own standup marks the expert-training design as "AT RISK" because
"We can't really choose the right experts to train without knowing how
our AAII proxies cluster and covary with true AAII" (<a
href="https://github.com/marin-community/marin/issues/9171">#9171</a>).
Whether experts run "as two separate expert runs, or ... one combined
curriculum, or both" is open (<a
href="https://discord.com/channels/1354881461060243556/1550202720433086504/1550234733495980173">Feuer</a>).</p>
<h2 id="points-worth-raising-with-your-colleague">Points worth raising
with your colleague</h2>
<ul>
<li>The coverage map and the curriculum use different vocabularies (151
competencies versus 1,999 capabilities), and Hall's TaskCompendium adds
a third tag ontology. No document maps them onto each other yet. Power
has said he wants to iterate before locking any taxonomy down (<a
data-orig-href="https://github.com/marin-community/marin/pull/9187#issuecomment-5690740119" href="https://github.com/marin-community/marin/pull/9187#issuecomment-5690740119">review</a>).</li>
<li>"Environment per element" is the intent of Muchane's pipeline,
scoped to the subset of gap competencies that are coding-shaped and have
practitioner repositories. Reviewers already flagged that this misses
law, medicine, and office work, which is where several of the named gaps
sit.</li>
<li>GLM-5.3's measured strengths so far are extraction, classification,
and rollouts. Its one-attempt pass rate of 9 percent on RST tasks means
teacher rollouts alone are a thin source for hard agentic tasks.</li>
<li>Everything here is dated September 22, 2026, and moving weekly. The
curriculum PRs merged on September 20, the routing release is dated
September 18, and the pipeline doc is a living document with open
comments.</li>
</ul>
</div>
<footer class="provenance">
<p><em>Corpus: 2026-09-22 ~05:30 UTC (marinmirror 221,146 chunks; weekly summaries through 2026-09-14_2026-09-20) &middot; Generated: 2026-09-22 06:03 UTC &middot; Viewing: <span id="viewing-time"></span></em></p>
<blockquote>
<p><em>Data: marinmirror, 221,146 chunks, built 1h ago (refreshed this
session, 2026-09-22) · summaries through 2026-09-14_2026-09-20.
Supplemented with the public GCS coverage report, the public Marina
applet catalog JSON (revision 9), Mark Muchane's public design doc and
its four linked pastebins, the task-curriculum README, workflow, and
HISTORY files on marin main, and GitHub PR metadata via the gh CLI. The
Coverage-Directed Agentic Environment Generation, Unified Task Format,
Environment Execution, and Expert Training design docs are
access-restricted and were not read.</em></p>
<p><em>Query: "I need to share with a business colleague more
information about how Marin is generating RL environments to cover their
taxonomy and curriculum. What does it mean to cover this taxonomy and
curriculum, and specifically how are we using GLM-5.3 to generate an
environment for each element of this taxonomy/curriculum?"</em></p>
<p><em>Sub-queries: "TaskTrove competency coverage audit #8879 and
follow-ups" · "cross-domain task curriculum #9232/#9294 design and
build" · "Mark Muchane GLM 5.3 synthetic task generation and the RST
rollouts thread" · "task environment contract #9122 and TaskCompendium
#9187" · "TaskTrove to Harbor conversion and SkyRL selection" ·
"curriculum RL launcher and RL consumption of environments" ·
"post-training data plan and OT-Next breakout" · "GLM-5.3 uses, serving,
and rationale" · "domain-specific task generation efforts" · "weekly
summaries Aug 31 to Sep 20" · "code lane: task_curriculum,
taskcompendium, tasktrove, curriculum_rl"</em></p>
</blockquote>
</footer>
</div>
</body>
</html>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment