Skip to content

Instantly share code, notes, and snippets.

@hammer
Last active September 29, 2026 08:01
Show Gist options
  • Select an option

  • Save hammer/a17688f9fcce5879ea7da426940ce2c4 to your computer and use it in GitHub Desktop.

Select an option

Save hammer/a17688f9fcce5879ea7da426940ce2c4 to your computer and use it in GitHub Desktop.
IFM vs. Marin: pretraining, data and post-training stacks
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>IFM vs. Marin: pretraining, data and post-training stacks</title>
<style>
:root {
--ink: #1a1a2e; --ink-secondary: #555770; --ink-faint: #8a8a9a;
--accent: #8b2500; --accent-light: #c4530a;
--surface: #faf9f6; --surface-raised: #f0eeea; --surface-code: #f5f3ef;
--rule: #d4d0c8; --link: #8b2500; --link-hover: #c4530a; --ref-bg: #f7f5f1;
--content-width: 650px; --sidenote-width: 230px; --sidenote-gap: 30px;
}
@media (prefers-color-scheme: dark) {
:root:not([data-theme="light"]) {
--ink: #d8d5cf; --ink-secondary: #9e9bab; --ink-faint: #6e6b7b;
--accent: #d4764e; --accent-light: #e8956e;
--surface: #1a1a24; --surface-raised: #242430; --surface-code: #20202c;
--rule: #33333f; --link: #d4764e; --link-hover: #e8956e; --ref-bg: #1e1e2a;
}
}
:root[data-theme="dark"] {
--ink: #d8d5cf; --ink-secondary: #9e9bab; --ink-faint: #6e6b7b;
--accent: #d4764e; --accent-light: #e8956e;
--surface: #1a1a24; --surface-raised: #242430; --surface-code: #20202c;
--rule: #33333f; --link: #d4764e; --link-hover: #e8956e; --ref-bg: #1e1e2a;
}
* { margin: 0; padding: 0; box-sizing: border-box; }
body {
background: var(--surface); color: var(--ink);
font-family: 'Helvetica Neue', Helvetica, Arial, sans-serif;
font-size: 16px; line-height: 1.7;
-webkit-font-smoothing: antialiased;
}
.page {
max-width: calc(var(--content-width) + var(--sidenote-width) + var(--sidenote-gap) + 80px);
margin: 0 auto; padding: 3rem 40px 4rem; position: relative;
}
@media (max-width: 1060px) {
.page { max-width: 100%; padding: 2rem 1.5rem 3rem; }
}
.content { max-width: var(--content-width); }
.paper-header { max-width: var(--content-width); margin-bottom: 2.5rem; padding-bottom: 2rem; }
.paper-header h1 {
font-family: 'Palatino Linotype', Palatino, 'Book Antiqua', Georgia, serif;
font-size: 2rem; font-weight: 600; line-height: 1.25; color: var(--ink);
text-wrap: balance; margin-bottom: 0.6rem; letter-spacing: -0.01em;
}
.paper-meta { font-size: 0.875rem; color: var(--ink-secondary); line-height: 1.5; }
.paper-meta .author { font-weight: 500; }
.paper-meta .mumwelt-link { color: var(--ink-secondary); text-decoration: none; border-bottom: 1px dotted var(--ink-faint); }
.paper-meta .mumwelt-link:hover { color: var(--accent); border-bottom-color: var(--accent); }
.prompt-box {
max-width: var(--content-width); margin-bottom: 2.5rem; position: relative;
}
.prompt-label {
position: absolute; left: -0.6em; top: 50%; transform: translateY(-50%);
font-family: 'Palatino Linotype', Palatino, 'Book Antiqua', Georgia, serif;
font-size: 5rem; font-weight: 700; color: var(--ink); opacity: 0.08;
line-height: 1; pointer-events: none; user-select: none;
}
.prompt-text {
font-family: 'Palatino Linotype', Palatino, 'Book Antiqua', Georgia, serif;
font-style: italic; font-size: 1.15rem; line-height: 1.55; color: var(--ink);
position: relative;
}
.abstract { margin-bottom: 2.5rem; max-width: var(--content-width); }
.abstract-label {
font-size: 0.7rem; font-weight: 600; text-transform: uppercase;
letter-spacing: 0.1em; color: var(--ink-secondary); margin-bottom: 0.5rem;
}
.abstract p { font-size: 0.92rem; line-height: 1.75; color: var(--ink); }
h1, h2, h3 {
font-family: 'Palatino Linotype', Palatino, 'Book Antiqua', Georgia, serif;
font-weight: 600; color: var(--ink); text-wrap: balance;
}
h2 { font-size: 1.4rem; margin-top: 2.5rem; margin-bottom: 0.75rem; letter-spacing: -0.005em; }
h3 { font-size: 1.1rem; margin-top: 1.75rem; margin-bottom: 0.5rem; }
p { margin-bottom: 1rem; max-width: var(--content-width); }
a { color: var(--link); text-decoration: none; border-bottom: 1px solid transparent;
transition: border-color 0.15s, color 0.15s; }
a:hover { color: var(--link-hover); border-bottom-color: var(--link-hover); }
.date-label { cursor: default; border-bottom: 1px dotted var(--ink-faint); }
strong { font-weight: 600; }
.sidenote-checkbox { display: none; }
.sidenote-toggle { display: none; }
.sidenote {
float: right; clear: right; width: var(--sidenote-width);
margin-right: calc(-1 * (var(--sidenote-width) + var(--sidenote-gap)));
margin-top: 0.2rem; margin-bottom: 1rem;
font-size: 0.8rem; line-height: 1.5; color: var(--ink-secondary);
}
.sidenote-number { font-size: 0.7rem; font-weight: 600; color: var(--accent); margin-right: 0.3em; }
@media (max-width: 1060px) {
.sidenote-toggle {
display: inline; cursor: pointer; color: var(--accent);
font-size: 0.78rem; font-weight: 600; user-select: none;
}
.sidenote {
float: none; display: none; width: 100%; margin: 0.4rem 0 0.75rem 0;
font-size: 0.84rem; padding: 0.6rem 0.9rem; background: var(--surface-raised);
border-radius: 4px; border-left: 2px solid var(--accent);
}
.sidenote-checkbox:checked + .sidenote { display: block; }
.sidenote-number { display: none; }
}
blockquote { border-left: 2px solid var(--rule); padding-left: 1.25rem;
margin: 1.25rem 0; color: var(--ink-secondary); font-style: italic; }
code { font-family: 'SF Mono', Menlo, Consolas, monospace; font-size: 0.85em;
background: var(--surface-code); padding: 0.15em 0.35em; border-radius: 3px; }
pre { background: var(--surface-code); border: 1px solid var(--rule); border-radius: 4px;
padding: 1rem 1.25rem; overflow-x: auto; margin: 1.25rem 0; max-width: var(--content-width); }
pre code { background: none; padding: 0; font-size: 0.82rem; line-height: 1.6; }
ul, ol { margin-bottom: 1rem; padding-left: 1.5rem; max-width: var(--content-width); }
li { margin-bottom: 0.35rem; }
li::marker { color: var(--ink-faint); }
table { max-width: var(--content-width); border-collapse: collapse; width: 100%;
margin: 1.25rem 0; font-variant-numeric: tabular-nums; font-size: 0.9rem; }
thead { border-top: 2px solid var(--ink); border-bottom: 1px solid var(--ink); }
th { font-weight: 600; padding: 0.3rem 0.75rem 0.35rem; text-align: left;
line-height: 1.2; white-space: nowrap; font-size: 0.82rem; vertical-align: bottom; }
td { padding: 0.35rem 0.75rem; border: none; vertical-align: top; }
tbody { border-bottom: 1.5px solid var(--ink); }
th:first-child, td:first-child { padding-left: 0; }
th:last-child, td:last-child { padding-right: 0; }
figure { margin: 2rem 0; max-width: var(--content-width); }
figure img { width: 100%; border-radius: 3px; border: 1px solid var(--rule); }
figcaption { font-size: 0.8rem; color: var(--ink-secondary); margin-top: 0.5rem;
line-height: 1.5; font-style: italic; }
footer.provenance {
margin-top: 3rem; padding-top: 1rem; max-width: var(--content-width);
font-size: 0.78rem; line-height: 1.6; color: var(--ink-faint);
}
footer.provenance p { margin-bottom: 0.3rem; }
footer.provenance a { color: var(--ink-faint); }
footer.provenance blockquote { margin: 0; padding: 0; border: none; color: inherit; }
a[data-hover-title] { position: relative; }
.hover-card {
position: absolute; bottom: 100%; left: 50%; transform: translateX(-50%);
width: 320px; max-width: 90vw; padding: 0.65rem 0.8rem;
background: var(--surface-raised); border: 1px solid var(--rule);
border-radius: 6px; box-shadow: 0 4px 12px rgba(0,0,0,0.1);
font-size: 0.78rem; line-height: 1.45; color: var(--ink);
pointer-events: none; z-index: 100; margin-bottom: 6px;
opacity: 0; transition: opacity 0.12s;
}
a[data-hover-title]:hover .hover-card,
a[data-hover-title]:focus .hover-card,
a.cite:hover .hover-card { opacity: 1; }
.hover-card .hc-title { font-weight: 600; margin-bottom: 0.2rem; }
.hover-card .hc-meta { font-size: 0.72rem; color: var(--ink-faint); margin-bottom: 0.25rem; }
.hover-card .hc-status {
display: inline-block; font-size: 0.65rem; font-weight: 600;
text-transform: uppercase; letter-spacing: 0.04em;
padding: 0.1em 0.4em; border-radius: 3px; margin-right: 0.4em;
}
.hc-status-open { background: #fff3e0; color: #e65100; }
.hc-status-closed, .hc-status-merged { background: #e8f5e9; color: #2e7d32; }
@media (prefers-color-scheme: dark) {
:root:not([data-theme="light"]) .hover-card { box-shadow: 0 4px 12px rgba(0,0,0,0.3); }
:root:not([data-theme="light"]) .hc-status-open { background: #3a2a10; color: #ffb74d; }
:root:not([data-theme="light"]) .hc-status-closed,
:root:not([data-theme="light"]) .hc-status-merged { background: #1b3a1e; color: #66bb6a; }
}
:root[data-theme="dark"] .hover-card { box-shadow: 0 4px 12px rgba(0,0,0,0.3); }
:root[data-theme="dark"] .hc-status-open { background: #3a2a10; color: #ffb74d; }
:root[data-theme="dark"] .hc-status-closed,
:root[data-theme="dark"] .hc-status-merged { background: #1b3a1e; color: #66bb6a; }
.hover-card .hc-desc { color: var(--ink-secondary); }
.cite {
font-size: 0.72rem; vertical-align: super; line-height: 0;
color: var(--accent); font-weight: 600; text-decoration: none;
border-bottom: none !important; position: relative;
}
.cite:hover { color: var(--link-hover); }
.references { margin-top: 3rem; max-width: var(--content-width); }
.references h2 { font-size: 1.15rem; margin-bottom: 1rem; }
.ref-list { list-style: none; padding: 0; counter-reset: ref; }
.ref-list li {
counter-increment: ref; display: flex; align-items: baseline;
gap: 0.5em; font-size: 0.82rem; line-height: 1.55;
margin-bottom: 0.4rem; color: var(--ink-secondary);
}
.ref-list li::before {
content: "[" counter(ref) "]"; flex-shrink: 0;
font-variant-numeric: tabular-nums; color: var(--ink-faint);
font-size: 0.78rem; min-width: 2.2em;
}
.ref-list .ref-body { flex: 1; min-width: 0; }
.ref-list .ref-title { font-weight: 500; color: var(--ink); }
.ref-list .ref-url {
font-family: 'SF Mono', Menlo, Consolas, monospace; font-size: 0.75rem;
color: var(--ink-faint); word-break: break-all; margin-left: 0.4em;
}
.ref-list .ref-url a { color: var(--ink-faint); border-bottom: none; }
.ref-list .ref-url a:hover { color: var(--link-hover); }
.ref-back {
color: var(--accent); text-decoration: none; border-bottom: none !important;
margin-left: 0.3em; font-size: 0.78rem;
}
.ref-back:hover { color: var(--link-hover); }
.status {
display: inline-block; font-size: 0.65rem; font-weight: 600;
text-transform: uppercase; letter-spacing: 0.04em;
padding: 0.15em 0.5em; border-radius: 3px; vertical-align: middle;
}
.status-done { background: #e8f5e9; color: #2e7d32; }
.status-open { background: #fff3e0; color: #e65100; }
.status-blocked { background: #fce4ec; color: #c62828; }
@media (prefers-color-scheme: dark) {
:root:not([data-theme="light"]) .status-done { background: #1b3a1e; color: #66bb6a; }
:root:not([data-theme="light"]) .status-open { background: #3a2a10; color: #ffb74d; }
:root:not([data-theme="light"]) .status-blocked { background: #3a1520; color: #ef9a9a; }
}
:root[data-theme="dark"] .status-done { background: #1b3a1e; color: #66bb6a; }
:root[data-theme="dark"] .status-open { background: #3a2a10; color: #ffb74d; }
:root[data-theme="dark"] .status-blocked { background: #3a1520; color: #ef9a9a; }
.cite-card {
position: absolute; bottom: 100%; left: 50%; transform: translateX(-50%);
width: 360px; max-width: 90vw; padding: 0.65rem 0.8rem;
background: var(--surface-raised); border: 1px solid var(--rule);
border-radius: 6px; box-shadow: 0 4px 12px rgba(0,0,0,0.1);
font-size: 0.78rem; line-height: 1.45; color: var(--ink);
pointer-events: none; z-index: 100; margin-bottom: 6px;
opacity: 0; transition: opacity 0.12s;
font-weight: 400; vertical-align: baseline; text-align: left;
}
a.cite:hover .cite-card { opacity: 1; }
@media (prefers-color-scheme: dark) {
:root:not([data-theme="light"]) .cite-card { box-shadow: 0 4px 12px rgba(0,0,0,0.3); }
}
:root[data-theme="dark"] .cite-card { box-shadow: 0 4px 12px rgba(0,0,0,0.3); }
@media print {
body { font-size: 11pt; }
.page { max-width: 100%; padding: 0; }
.sidenote { float: right; width: 180px; margin-right: -210px; }
.sidenote-toggle { display: none; }
a { color: inherit; border-bottom: none; }
.hover-card { display: none; }
.cite-card { display: none; }
}
.katex{font:normal 1.21em KaTeX_Main,Times New Roman,serif;line-height:1.2;position:relative;text-indent:0;text-rendering:auto}.katex *{-ms-high-contrast-adjust:none!important;border-color:currentColor}.katex .katex-version:after{content:"0.18.9"}.katex .katex-mathml{border:0;-webkit-clip-path:inset(50%);clip-path:inset(50%);height:1px;overflow:hidden;padding:0;position:absolute;width:1px}.katex .katex-html>.katex-newline{display:block}.katex .katex-base{position:relative;white-space:nowrap;width:-webkit-min-content;width:-moz-min-content;width:min-content}.katex .katex-base,.katex .katex-strut{display:inline-block}.katex .textbf{font-weight:700}.katex .textit{font-style:italic}.katex .textrm{font-family:KaTeX_Main}.katex .textsf{font-family:KaTeX_SansSerif}.katex .texttt{font-family:KaTeX_Typewriter}.katex .mathnormal{font-family:KaTeX_Math;font-style:italic}.katex .mathit{font-family:KaTeX_Main;font-style:italic}.katex .mathrm{font-style:normal}.katex .mathbf{font-family:KaTeX_Main;font-weight:700}.katex .boldsymbol{font-family:KaTeX_Math;font-style:italic;font-weight:700}.katex .amsrm,.katex .mathbb,.katex .textbb{font-family:KaTeX_AMS}.katex .mathcal{font-family:KaTeX_Caligraphic}.katex .mathfrak,.katex .textfrak{font-family:KaTeX_Fraktur}.katex .mathboldfrak,.katex .textboldfrak{font-family:KaTeX_Fraktur;font-weight:700}.katex .mathtt{font-family:KaTeX_Typewriter}.katex .mathscr,.katex .textscr{font-family:KaTeX_Script}.katex .mathsf,.katex .textsf{font-family:KaTeX_SansSerif}.katex .mathboldsf,.katex .textboldsf{font-family:KaTeX_SansSerif;font-weight:700}.katex .mathitsf,.katex .mathsfit,.katex .textitsf{font-family:KaTeX_SansSerif;font-style:italic}.katex .mainrm{font-family:KaTeX_Main;font-style:normal}.katex .vlist-t{border-collapse:collapse;display:inline-table;table-layout:fixed}.katex .vlist-r{display:table-row}.katex .vlist{display:table-cell;position:relative;vertical-align:bottom}.katex .vlist>span{display:block;height:0;position:relative}.katex .vlist>span>span{display:inline-block}.katex .vlist>span>.pstrut{overflow:hidden;width:0}.katex .vlist-t2{margin-right:-2px}.katex .vlist-s{display:table-cell;font-size:1px;min-width:2px;vertical-align:bottom;width:2px}.katex .katex-vbox{align-items:baseline;display:inline-flex;flex-direction:column}.katex .katex-thinbox{display:inline-flex;flex-direction:row;max-width:0;width:0}.katex .msupsub{text-align:left}.katex .mfrac>span>span{text-align:center}.katex .mfrac .frac-line{border-bottom-style:solid;display:inline-block;width:100%}.katex .katex-hdashline,.katex .katex-hline,.katex .katex-overline .overline-line,.katex .katex-rule,.katex .katex-underline .underline-line,.katex .mfrac .frac-line{min-height:1px}.katex .mspace{display:inline-block}.katex .katex-smash{display:inline;line-height:0}.katex .clap,.katex .llap,.katex .rlap{position:relative;width:0}.katex .clap>.katex-inner,.katex .llap>.katex-inner,.katex .rlap>.katex-inner{position:absolute}.katex .clap>.katex-fix,.katex .llap>.katex-fix,.katex .rlap>.katex-fix{display:inline-block}.katex .llap>.katex-inner{right:0}.katex .clap>.katex-inner,.katex .rlap>.katex-inner{left:0}.katex .clap>.katex-inner>span{margin-left:-50%;margin-right:50%}.katex .katex-rule{border:0 solid;display:inline-block;position:relative}.katex .katex-hline,.katex .katex-overline .overline-line,.katex .katex-underline .underline-line{border-bottom-style:solid;display:inline-block;width:100%}.katex .katex-hdashline{border-bottom-style:dashed;display:inline-block;width:100%}.katex .sqrt>.katex-root{margin-left:.2777777778em;margin-right:-.5555555556em}.katex .fontsize-ensurer.reset-size1.size1,.katex .katex-sizing.reset-size1.size1{font-size:1em}.katex .fontsize-ensurer.reset-size1.size2,.katex .katex-sizing.reset-size1.size2{font-size:1.2em}.katex .fontsize-ensurer.reset-size1.size3,.katex .katex-sizing.reset-size1.size3{font-size:1.4em}.katex .fontsize-ensurer.reset-size1.size4,.katex .katex-sizing.reset-size1.size4{font-size:1.6em}.katex .fontsize-ensurer.reset-size1.size5,.katex .katex-sizing.reset-size1.size5{font-size:1.8em}.katex .fontsize-ensurer.reset-size1.size6,.katex .katex-sizing.reset-size1.size6{font-size:2em}.katex .fontsize-ensurer.reset-size1.size7,.katex .katex-sizing.reset-size1.size7{font-size:2.4em}.katex .fontsize-ensurer.reset-size1.size8,.katex .katex-sizing.reset-size1.size8{font-size:2.88em}.katex .fontsize-ensurer.reset-size1.size9,.katex .katex-sizing.reset-size1.size9{font-size:3.456em}.katex .fontsize-ensurer.reset-size1.size10,.katex .katex-sizing.reset-size1.size10{font-size:4.148em}.katex .fontsize-ensurer.reset-size1.size11,.katex .katex-sizing.reset-size1.size11{font-size:4.976em}.katex .fontsize-ensurer.reset-size2.size1,.katex .katex-sizing.reset-size2.size1{font-size:.8333333333em}.katex .fontsize-ensurer.reset-size2.size2,.katex .katex-sizing.reset-size2.size2{font-size:1em}.katex .fontsize-ensurer.reset-size2.size3,.katex .katex-sizing.reset-size2.size3{font-size:1.1666666667em}.katex .fontsize-ensurer.reset-size2.size4,.katex .katex-sizing.reset-size2.size4{font-size:1.3333333333em}.katex .fontsize-ensurer.reset-size2.size5,.katex .katex-sizing.reset-size2.size5{font-size:1.5em}.katex .fontsize-ensurer.reset-size2.size6,.katex .katex-sizing.reset-size2.size6{font-size:1.6666666667em}.katex .fontsize-ensurer.reset-size2.size7,.katex .katex-sizing.reset-size2.size7{font-size:2em}.katex .fontsize-ensurer.reset-size2.size8,.katex .katex-sizing.reset-size2.size8{font-size:2.4em}.katex .fontsize-ensurer.reset-size2.size9,.katex .katex-sizing.reset-size2.size9{font-size:2.88em}.katex .fontsize-ensurer.reset-size2.size10,.katex .katex-sizing.reset-size2.size10{font-size:3.4566666667em}.katex .fontsize-ensurer.reset-size2.size11,.katex .katex-sizing.reset-size2.size11{font-size:4.1466666667em}.katex .fontsize-ensurer.reset-size3.size1,.katex .katex-sizing.reset-size3.size1{font-size:.7142857143em}.katex .fontsize-ensurer.reset-size3.size2,.katex .katex-sizing.reset-size3.size2{font-size:.8571428571em}.katex .fontsize-ensurer.reset-size3.size3,.katex .katex-sizing.reset-size3.size3{font-size:1em}.katex .fontsize-ensurer.reset-size3.size4,.katex .katex-sizing.reset-size3.size4{font-size:1.1428571429em}.katex .fontsize-ensurer.reset-size3.size5,.katex .katex-sizing.reset-size3.size5{font-size:1.2857142857em}.katex .fontsize-ensurer.reset-size3.size6,.katex .katex-sizing.reset-size3.size6{font-size:1.4285714286em}.katex .fontsize-ensurer.reset-size3.size7,.katex .katex-sizing.reset-size3.size7{font-size:1.7142857143em}.katex .fontsize-ensurer.reset-size3.size8,.katex .katex-sizing.reset-size3.size8{font-size:2.0571428571em}.katex .fontsize-ensurer.reset-size3.size9,.katex .katex-sizing.reset-size3.size9{font-size:2.4685714286em}.katex .fontsize-ensurer.reset-size3.size10,.katex .katex-sizing.reset-size3.size10{font-size:2.9628571429em}.katex .fontsize-ensurer.reset-size3.size11,.katex .katex-sizing.reset-size3.size11{font-size:3.5542857143em}.katex .fontsize-ensurer.reset-size4.size1,.katex .katex-sizing.reset-size4.size1{font-size:.625em}.katex .fontsize-ensurer.reset-size4.size2,.katex .katex-sizing.reset-size4.size2{font-size:.75em}.katex .fontsize-ensurer.reset-size4.size3,.katex .katex-sizing.reset-size4.size3{font-size:.875em}.katex .fontsize-ensurer.reset-size4.size4,.katex .katex-sizing.reset-size4.size4{font-size:1em}.katex .fontsize-ensurer.reset-size4.size5,.katex .katex-sizing.reset-size4.size5{font-size:1.125em}.katex .fontsize-ensurer.reset-size4.size6,.katex .katex-sizing.reset-size4.size6{font-size:1.25em}.katex .fontsize-ensurer.reset-size4.size7,.katex .katex-sizing.reset-size4.size7{font-size:1.5em}.katex .fontsize-ensurer.reset-size4.size8,.katex .katex-sizing.reset-size4.size8{font-size:1.8em}.katex .fontsize-ensurer.reset-size4.size9,.katex .katex-sizing.reset-size4.size9{font-size:2.16em}.katex .fontsize-ensurer.reset-size4.size10,.katex .katex-sizing.reset-size4.size10{font-size:2.5925em}.katex .fontsize-ensurer.reset-size4.size11,.katex .katex-sizing.reset-size4.size11{font-size:3.11em}.katex .fontsize-ensurer.reset-size5.size1,.katex .katex-sizing.reset-size5.size1{font-size:.5555555556em}.katex .fontsize-ensurer.reset-size5.size2,.katex .katex-sizing.reset-size5.size2{font-size:.6666666667em}.katex .fontsize-ensurer.reset-size5.size3,.katex .katex-sizing.reset-size5.size3{font-size:.7777777778em}.katex .fontsize-ensurer.reset-size5.size4,.katex .katex-sizing.reset-size5.size4{font-size:.8888888889em}.katex .fontsize-ensurer.reset-size5.size5,.katex .katex-sizing.reset-size5.size5{font-size:1em}.katex .fontsize-ensurer.reset-size5.size6,.katex .katex-sizing.reset-size5.size6{font-size:1.1111111111em}.katex .fontsize-ensurer.reset-size5.size7,.katex .katex-sizing.reset-size5.size7{font-size:1.3333333333em}.katex .fontsize-ensurer.reset-size5.size8,.katex .katex-sizing.reset-size5.size8{font-size:1.6em}.katex .fontsize-ensurer.reset-size5.size9,.katex .katex-sizing.reset-size5.size9{font-size:1.92em}.katex .fontsize-ensurer.reset-size5.size10,.katex .katex-sizing.reset-size5.size10{font-size:2.3044444444em}.katex .fontsize-ensurer.reset-size5.size11,.katex .katex-sizing.reset-size5.size11{font-size:2.7644444444em}.katex .fontsize-ensurer.reset-size6.size1,.katex .katex-sizing.reset-size6.size1{font-size:.5em}.katex .fontsize-ensurer.reset-size6.size2,.katex .katex-sizing.reset-size6.size2{font-size:.6em}.katex .fontsize-ensurer.reset-size6.size3,.katex .katex-sizing.reset-size6.size3{font-size:.7em}.katex .fontsize-ensurer.reset-size6.size4,.katex .katex-sizing.reset-size6.size4{font-size:.8em}.katex .fontsize-ensurer.reset-size6.size5,.katex .katex-sizing.reset-size6.size5{font-size:.9em}.katex .fontsize-ensurer.reset-size6.size6,.katex .katex-sizing.reset-size6.size6{font-size:1em}.katex .fontsize-ensurer.reset-size6.size7,.katex .katex-sizing.reset-size6.size7{font-size:1.2em}.katex .fontsize-ensurer.reset-size6.size8,.katex .katex-sizing.reset-size6.size8{font-size:1.44em}.katex .fontsize-ensurer.reset-size6.size9,.katex .katex-sizing.reset-size6.size9{font-size:1.728em}.katex .fontsize-ensurer.reset-size6.size10,.katex .katex-sizing.reset-size6.size10{font-size:2.074em}.katex .fontsize-ensurer.reset-size6.size11,.katex .katex-sizing.reset-size6.size11{font-size:2.488em}.katex .fontsize-ensurer.reset-size7.size1,.katex .katex-sizing.reset-size7.size1{font-size:.4166666667em}.katex .fontsize-ensurer.reset-size7.size2,.katex .katex-sizing.reset-size7.size2{font-size:.5em}.katex .fontsize-ensurer.reset-size7.size3,.katex .katex-sizing.reset-size7.size3{font-size:.5833333333em}.katex .fontsize-ensurer.reset-size7.size4,.katex .katex-sizing.reset-size7.size4{font-size:.6666666667em}.katex .fontsize-ensurer.reset-size7.size5,.katex .katex-sizing.reset-size7.size5{font-size:.75em}.katex .fontsize-ensurer.reset-size7.size6,.katex .katex-sizing.reset-size7.size6{font-size:.8333333333em}.katex .fontsize-ensurer.reset-size7.size7,.katex .katex-sizing.reset-size7.size7{font-size:1em}.katex .fontsize-ensurer.reset-size7.size8,.katex .katex-sizing.reset-size7.size8{font-size:1.2em}.katex .fontsize-ensurer.reset-size7.size9,.katex .katex-sizing.reset-size7.size9{font-size:1.44em}.katex .fontsize-ensurer.reset-size7.size10,.katex .katex-sizing.reset-size7.size10{font-size:1.7283333333em}.katex .fontsize-ensurer.reset-size7.size11,.katex .katex-sizing.reset-size7.size11{font-size:2.0733333333em}.katex .fontsize-ensurer.reset-size8.size1,.katex .katex-sizing.reset-size8.size1{font-size:.3472222222em}.katex .fontsize-ensurer.reset-size8.size2,.katex .katex-sizing.reset-size8.size2{font-size:.4166666667em}.katex .fontsize-ensurer.reset-size8.size3,.katex .katex-sizing.reset-size8.size3{font-size:.4861111111em}.katex .fontsize-ensurer.reset-size8.size4,.katex .katex-sizing.reset-size8.size4{font-size:.5555555556em}.katex .fontsize-ensurer.reset-size8.size5,.katex .katex-sizing.reset-size8.size5{font-size:.625em}.katex .fontsize-ensurer.reset-size8.size6,.katex .katex-sizing.reset-size8.size6{font-size:.6944444444em}.katex .fontsize-ensurer.reset-size8.size7,.katex .katex-sizing.reset-size8.size7{font-size:.8333333333em}.katex .fontsize-ensurer.reset-size8.size8,.katex .katex-sizing.reset-size8.size8{font-size:1em}.katex .fontsize-ensurer.reset-size8.size9,.katex .katex-sizing.reset-size8.size9{font-size:1.2em}.katex .fontsize-ensurer.reset-size8.size10,.katex .katex-sizing.reset-size8.size10{font-size:1.4402777778em}.katex .fontsize-ensurer.reset-size8.size11,.katex .katex-sizing.reset-size8.size11{font-size:1.7277777778em}.katex .fontsize-ensurer.reset-size9.size1,.katex .katex-sizing.reset-size9.size1{font-size:.2893518519em}.katex .fontsize-ensurer.reset-size9.size2,.katex .katex-sizing.reset-size9.size2{font-size:.3472222222em}.katex .fontsize-ensurer.reset-size9.size3,.katex .katex-sizing.reset-size9.size3{font-size:.4050925926em}.katex .fontsize-ensurer.reset-size9.size4,.katex .katex-sizing.reset-size9.size4{font-size:.462962963em}.katex .fontsize-ensurer.reset-size9.size5,.katex .katex-sizing.reset-size9.size5{font-size:.5208333333em}.katex .fontsize-ensurer.reset-size9.size6,.katex .katex-sizing.reset-size9.size6{font-size:.5787037037em}.katex .fontsize-ensurer.reset-size9.size7,.katex .katex-sizing.reset-size9.size7{font-size:.6944444444em}.katex .fontsize-ensurer.reset-size9.size8,.katex .katex-sizing.reset-size9.size8{font-size:.8333333333em}.katex .fontsize-ensurer.reset-size9.size9,.katex .katex-sizing.reset-size9.size9{font-size:1em}.katex .fontsize-ensurer.reset-size9.size10,.katex .katex-sizing.reset-size9.size10{font-size:1.2002314815em}.katex .fontsize-ensurer.reset-size9.size11,.katex .katex-sizing.reset-size9.size11{font-size:1.4398148148em}.katex .fontsize-ensurer.reset-size10.size1,.katex .katex-sizing.reset-size10.size1{font-size:.2410800386em}.katex .fontsize-ensurer.reset-size10.size2,.katex .katex-sizing.reset-size10.size2{font-size:.2892960463em}.katex .fontsize-ensurer.reset-size10.size3,.katex .katex-sizing.reset-size10.size3{font-size:.337512054em}.katex .fontsize-ensurer.reset-size10.size4,.katex .katex-sizing.reset-size10.size4{font-size:.3857280617em}.katex .fontsize-ensurer.reset-size10.size5,.katex .katex-sizing.reset-size10.size5{font-size:.4339440694em}.katex .fontsize-ensurer.reset-size10.size6,.katex .katex-sizing.reset-size10.size6{font-size:.4821600771em}.katex .fontsize-ensurer.reset-size10.size7,.katex .katex-sizing.reset-size10.size7{font-size:.5785920926em}.katex .fontsize-ensurer.reset-size10.size8,.katex .katex-sizing.reset-size10.size8{font-size:.6943105111em}.katex .fontsize-ensurer.reset-size10.size9,.katex .katex-sizing.reset-size10.size9{font-size:.8331726133em}.katex .fontsize-ensurer.reset-size10.size10,.katex .katex-sizing.reset-size10.size10{font-size:1em}.katex .fontsize-ensurer.reset-size10.size11,.katex .katex-sizing.reset-size10.size11{font-size:1.1996142719em}.katex .fontsize-ensurer.reset-size11.size1,.katex .katex-sizing.reset-size11.size1{font-size:.2009646302em}.katex .fontsize-ensurer.reset-size11.size2,.katex .katex-sizing.reset-size11.size2{font-size:.2411575563em}.katex .fontsize-ensurer.reset-size11.size3,.katex .katex-sizing.reset-size11.size3{font-size:.2813504823em}.katex .fontsize-ensurer.reset-size11.size4,.katex .katex-sizing.reset-size11.size4{font-size:.3215434084em}.katex .fontsize-ensurer.reset-size11.size5,.katex .katex-sizing.reset-size11.size5{font-size:.3617363344em}.katex .fontsize-ensurer.reset-size11.size6,.katex .katex-sizing.reset-size11.size6{font-size:.4019292605em}.katex .fontsize-ensurer.reset-size11.size7,.katex .katex-sizing.reset-size11.size7{font-size:.4823151125em}.katex .fontsize-ensurer.reset-size11.size8,.katex .katex-sizing.reset-size11.size8{font-size:.578778135em}.katex .fontsize-ensurer.reset-size11.size9,.katex .katex-sizing.reset-size11.size9{font-size:.6945337621em}.katex .fontsize-ensurer.reset-size11.size10,.katex .katex-sizing.reset-size11.size10{font-size:.8336012862em}.katex .fontsize-ensurer.reset-size11.size11,.katex .katex-sizing.reset-size11.size11{font-size:1em}.katex .delimsizing.size1{font-family:KaTeX_Size1}.katex .delimsizing.size2{font-family:KaTeX_Size2}.katex .delimsizing.size3{font-family:KaTeX_Size3}.katex .delimsizing.size4{font-family:KaTeX_Size4}.katex .delimsizing.mult .delim-size1>span{font-family:KaTeX_Size1}.katex .delimsizing.mult .delim-size4>span{font-family:KaTeX_Size4}.katex .nulldelimiter{display:inline-block;width:.12em}.katex .delimcenter,.katex .op-symbol{position:relative}.katex .op-symbol.small-op{font-family:KaTeX_Size1}.katex .op-symbol.large-op{font-family:KaTeX_Size2}.katex .katex-accent>.vlist-t,.katex .op-limits>.vlist-t{text-align:center}.katex .katex-accent .accent-body{position:relative}.katex .katex-accent .accent-body:not(.accent-full){width:0}.katex .katex-overlay{display:block}.katex .mtable .vertical-separator{display:inline-block;min-width:1px}.katex .mtable .arraycolsep{display:inline-block}.katex .mtable .col-align-c>.vlist-t{text-align:center}.katex .mtable .col-align-l>.vlist-t{text-align:left}.katex .mtable .col-align-r>.vlist-t{text-align:right}.katex .svg-align{text-align:left}.katex svg{fill:currentColor;stroke:currentColor;display:block;height:inherit;position:absolute;width:100%}.katex svg path{stroke:none}.katex svg{fill-rule:nonzero;fill-opacity:1;stroke-width:1;stroke-linecap:butt;stroke-linejoin:miter;stroke-miterlimit:4;stroke-dasharray:none;stroke-dashoffset:0;stroke-opacity:1}.katex img{border-style:none;max-height:none;max-width:none;min-height:0;min-width:0}.katex .katex-stretchy{display:block;overflow:hidden;position:relative;width:100%}.katex .katex-stretchy:after,.katex .katex-stretchy:before{content:""}.katex .hide-tail{overflow:hidden;position:relative;width:100%}.katex .halfarrow-left{left:0;overflow:hidden;position:absolute;width:50.2%}.katex .halfarrow-right{overflow:hidden;position:absolute;right:0;width:50.2%}.katex .brace-left{left:0;overflow:hidden;position:absolute;width:25.1%}.katex .brace-center{left:25%;overflow:hidden;position:absolute;width:50%}.katex .brace-right{overflow:hidden;position:absolute;right:0;width:25.1%}.katex .x-arrow-pad{padding:0 .5em}.katex .cd-arrow-pad{padding:0 .55556em 0 .27778em}.katex .mover,.katex .munder,.katex .x-arrow{text-align:center}.katex .boxpad{padding:0 .3em}.katex .fbox,.katex .fcolorbox{border:.04em solid;box-sizing:border-box}.katex .cancel-pad{padding:0 .2em}.katex .cancel-lap{margin-left:-.2em;margin-right:-.2em}.katex .katex-sout{border-bottom-style:solid;border-bottom-width:.08em}.katex .angl{border-right:.049em solid;border-top:.049em solid;box-sizing:border-box;margin-right:.03889em}.katex .anglpad{padding:0 .03889em}.katex .reflectbox{display:inline-block;transform:scaleX(-1)}.katex .eqn-num:before{content:"(" counter(katexEqnNo) ")";counter-increment:katexEqnNo}.katex .mml-eqn-num:before{content:"(" counter(mmlEqnNo) ")";counter-increment:mmlEqnNo}.katex .mtr-glue{width:50%}.katex .cd-vert-arrow{display:inline-block;position:relative}.katex .cd-label-left{display:inline-block;position:absolute;right:calc(50% + .3em);text-align:left}.katex .cd-label-right{display:inline-block;left:calc(50% + .3em);position:absolute;text-align:right}.katex-display{display:block;margin:1em 0;text-align:center}.katex-display>.katex{display:block;text-align:center;white-space:nowrap}.katex-display>.katex>.katex-html{display:block;position:relative}.katex-display>.katex>.katex-html>.katex-tag{position:absolute;right:0}.katex-display.leqno>.katex>.katex-html>.katex-tag{left:0;right:auto}.katex-display.fleqn>.katex{padding-left:2em;text-align:left}body{counter-reset:katexEqnNo mmlEqnNo}
</style>
<script>(function() {
function __mumInit() {
// Restore external links that htmlpreview.github.io's loader rewrote to
// in-page anchors (it treats any href containing '#' as a local anchor).
// Runs now (htmlpreview re-executes this script after its rewrite), again
// after delays, and on click as a last line of defense.
function restoreOrigHrefs() {
document.querySelectorAll('a[data-orig-href]').forEach(function(a) {
var orig = a.getAttribute('data-orig-href');
if (orig && a.getAttribute('href') !== orig) a.setAttribute('href', orig);
});
}
restoreOrigHrefs();
setTimeout(restoreOrigHrefs, 100);
setTimeout(restoreOrigHrefs, 1000);
document.addEventListener('click', function(e) {
var t = e.target && e.target.closest ? e.target.closest('a[data-orig-href]') : null;
if (t) { var o = t.getAttribute('data-orig-href'); if (o && t.getAttribute('href') !== o) t.setAttribute('href', o); }
}, true);
var vt = document.getElementById('viewing-time');
if (vt) {
function pad(n) { return n < 10 ? '0' + n : n; }
function updateTime() {
var d = new Date();
vt.textContent = d.getUTCFullYear() + '-' + pad(d.getUTCMonth()+1) + '-' + pad(d.getUTCDate()) + ' ' + pad(d.getUTCHours()) + ':' + pad(d.getUTCMinutes()) + ' UTC';
}
updateTime();
setInterval(updateTime, 60000);
}
document.querySelectorAll('a[data-hover-title]').forEach(function(a) {
var card = document.createElement('span');
card.className = 'hover-card';
var title = a.getAttribute('data-hover-title') || '';
var status = a.getAttribute('data-hover-status') || '';
var owner = a.getAttribute('data-hover-owner') || '';
var desc = a.getAttribute('data-hover-desc') || '';
var statusCls = 'hc-status hc-status-' + status.toLowerCase().replace(/[^a-z]/g, '');
var html = '<div class="hc-title">' + title + '</div>';
var meta = [];
if (status) meta.push('<span class="' + statusCls + '">' + status + '</span>');
if (owner) meta.push(owner);
if (meta.length) html += '<div class="hc-meta">' + meta.join(' ') + '</div>';
if (desc) html += '<div class="hc-desc">' + desc + '</div>';
card.innerHTML = html;
a.appendChild(card);
});
document.querySelectorAll('a.cite').forEach(function(a) {
var href = a.getAttribute('href') || '';
if (!href.startsWith('#ref-')) return;
var li = document.getElementById(href.slice(1));
if (!li) return;
var body = li.querySelector('.ref-body');
if (!body) return;
var refUrl = body.querySelector('.ref-url a');
var refHref = refUrl ? refUrl.getAttribute('href') : '';
var hoverSource = refHref ? document.querySelector('a[data-hover-title][href="' + refHref + '"]') : null;
if (hoverSource) {
var card = document.createElement('span');
card.className = 'hover-card';
var t = hoverSource.getAttribute('data-hover-title') || '';
var s = hoverSource.getAttribute('data-hover-status') || '';
var o = hoverSource.getAttribute('data-hover-owner') || '';
var d = hoverSource.getAttribute('data-hover-desc') || '';
var sc = 'hc-status hc-status-' + s.toLowerCase().replace(/[^a-z]/g, '');
var h = '<div class="hc-title">' + t + '</div>';
var m = [];
if (s) m.push('<span class="' + sc + '">' + s + '</span>');
if (o) m.push(o);
if (m.length) h += '<div class="hc-meta">' + m.join(' ') + '</div>';
if (d) h += '<div class="hc-desc">' + d + '</div>';
card.innerHTML = h;
a.appendChild(card);
} else {
var title = body.querySelector('.ref-title');
if (!title) return;
var card = document.createElement('span');
card.className = 'cite-card';
card.textContent = title.textContent;
a.appendChild(card);
}
});
}
if (document.readyState !== 'loading') { __mumInit(); }
else { document.addEventListener('DOMContentLoaded', __mumInit); }
})();</script>
</head>
<body>
<div class="page">
<header class="paper-header">
<h1>IFM vs. Marin: pretraining, data and post-training stacks</h1>
<div class="paper-meta">Posed by <span class="author">hammer</span>, answered by <a class="mumwelt-link" href="https://github.com/marin-community/mumwelt">mumwelt</a> &middot; <span class="date-label" title="Generated 2026-09-29 08:01 UTC">Published 2026-09-29</span></div>
</header>
<div class="prompt-box">
<div class="prompt-label">?</div>
<div class="prompt-text">Do a detailed analysis of https://github.com/ifm-ai/xllm, a new library for specifying and training large LLMs, and compare it to what we have in Levanter. Consider flexibility in specifying new architectures, completeness of the linear and other scalable attention mechanisms implemented, FP8 and FP4 training, completeness of kernels for different GPUs, ability to train on non-GPUs, dimensions of parallelism implemented, ease of evolution by agents, code complexity, test coverage, performance, and any other measure important for a team building frontier LLMs that needs to choose between Levanter and xLLM. Then compare the data curation and mixing code in https://github.com/ifm-ai/pretraining-data-toolkit with Marin&#x27;s Datakit, and https://github.com/ifm-ai/RL360 with MarinSkyRL and other post-training code from Levanter and Marin, including what IFM&#x27;s talk at https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf says about data curation and mixing.</div>
</div>
<div class="content">
<p><em>Code read at IFM's xLLM <a
href="https://github.com/ifm-ai/xllm/tree/889db388d021c92164e7a675f6884dd40ea117c7"><code>889db38</code></a>,
xattn <a
href="https://github.com/ifm-ai/xattn/tree/0d6d73b37acc644d5559007fad370185b78ddbb5"><code>0d6d73b</code></a>,
xbridges <a
href="https://github.com/ifm-ai/xbridges/tree/227f3fd26e30b024ce09bc7c556c64616c13276e"><code>227f3fd</code></a>,
pretraining-data-toolkit <a
href="https://github.com/ifm-ai/pretraining-data-toolkit/tree/cd93e116a8217c9e9a45e104b0ea28ac653a998f"><code>cd93e11</code></a>
and RL360 <a
href="https://github.com/ifm-ai/RL360/tree/5b548b6948e6e78af99307a38ee96057ea571faa"><code>5b548b6</code></a>,
and at Marin <code>main</code> <a
href="https://github.com/marin-community/marin/tree/4aa26ec73d248b974766f8940a896a8c888d8269"><code>4aa26ec</code></a>
and MarinSkyRL <a
href="https://github.com/marin-community/MarinSkyRL/tree/0cdccc9229f47c8eadff90aa2033a5b4604266aa"><code>0cdccc9</code></a>,
all as of 2026-09-28 or 2026-09-29. Marin history comes from its GitHub,
Discord and W&amp;B records and its <a
href="https://openathena.ai/blog/marin-data-pipeline-overview/">Datakit
blog post</a>; IFM history from its blog, model and dataset cards,
public W&amp;B logs and a <a
href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf">talk
on its data work</a>. Nothing was run on a GPU or TPU: every speed below
was reported or logged by its authors.</em></p>
<h2 id="summary">Summary</h2>
<p><a href="https://github.com/ifm-ai/xllm">xLLM</a> is the PyTorch
framework that MBZUAI's Institute of Foundation Models (IFM) used to
pretrain its K2 Horizon models, which range from a 0.9B dense model to a
375B-A23B mixture-of-experts (MoE) model. IFM opened the repository on
2026-09-02 and pushed the code on 2026-09-28. <em>Levanter</em> here
means Marin's JAX stack: the Levanter library, the Haliax named-tensor
library, and Grug, the copy-and-edit model code now training Marin's
535B-A23B MoE model on 704 GB200 GPUs. Sections 13 and 14 extend the
comparison to data curation and post-training: IFM's
pretraining-data-toolkit and RL360 against Marin's Datakit and
MarinSkyRL.</p>
<p>Levanter is the stronger base for frontier pretraining. It has what
xLLM lacks: Muon-family optimizers, expert parallelism across a 64-GPU
NVLink domain, proven use on Blackwell GPUs, and TPU support. Neither
stack's production runs use pipeline parallelism, gradient accumulation
or FP8, but Levanter has working code for all three and xLLM's public
code has none. Levanter also tests far more and is built for agents to
change.</p>
<p>xLLM has trained dense models at scale. IFM's public logs show an
internal version of it pretraining a 7B model on 21.9T tokens at about
48.6% model FLOPs utilization (MFU) by its own formula, and extending
context to 512K tokens. It supports sparse MoE and probably pretrained
the 375B-A23B as well, though IFM has published no throughput for that
run. Above its kernels it tests only its data pipeline, it has no CI,
and its most original component, a long-context layer called Gekko,
appears in none of the released models.</p>
<p>Levanter has weaknesses too. Its GPU speed depends on a patched XLA
plugin and exactly pinned, fast-moving kernel packages, and the hero's
GB200 path rests mainly on one engineer. No GPU tests run on pull
requests. And Marin has no 7–8B dense pretraining measurement on GPUs to
set beside xLLM's numbers.</p>
<p>For data curation, Marin has the more complete code and IFM has
released far more data. IFM's public data toolkit covers the middle of
its pipeline: labeling documents with public classifiers, sorting them
into quality categories and shuffling them, in about 3,100 lines with no
tests. The Common Crawl pipeline behind K2 Horizon's web text is public
from IFM's earlier TxT360 work, but the processing of its newer sources,
most of its synthesis code and its mixture search are not. IFM has
released about 13.5 trillion tokens of training data, mostly synthetic.
Marin's Datakit runs from download through deduplication,
decontamination, topic and quality bucketing and tokenization, with
hundreds of tests. Marin rehosts none of its pretraining corpus, several
of whose sources forbid redistribution, and the code that chose its
hero's buckets and mixture weights is not on <code>main</code>. Both
labs found that searching for mixture weights gains less than adding new
or more diverse data.</p>
<p>For post-training, Marin has the more complete RL code and IFM the
results. Both run RL in PyTorch on Megatron and run agentic tasks
through Harbor. IFM's RL360 is a snapshot of five forked projects with
no tests of its own and one demonstration recipe that IFM says has not
been validated end to end; IFM trained its shipped experts with internal
versions of it and other harnesses, and its merge code is private.
MarinSkyRL, Marin's fork of SkyRL, has about 2,300 tests but changes
weekly: in the four days before the commit read here, it deleted its
FSDP trainer and replaced its training loop. Because Marin pretrains and
fine-tunes in JAX, its RL needs a PyTorch port of each model, a cost IFM
avoids. IFM has shipped post-trained models at six sizes; Marin's first
official post-trained model has been pushed back a week from its planned
October 6 debut.</p>
<p><strong>Recommendation.</strong> Marin should stay on Levanter, run
one same-hardware dense benchmark against xLLM, count causal attention
in its MFU, and study IFM's long-context schedule, Gekko and MoVA. From
the data and post-training comparison, it should test IFM's released
datasets, favor new data over further mixture search, move its hero's
data and SFT code onto <code>main</code>, and compare weight merging
with distillation. For other teams, the hardware decides the pretraining
stack: on TPUs or GB200 racks, choose Levanter; on Hopper clusters under
Slurm, for dense models or long-context research, xLLM is a credible
start; for frontier MoE on Hopper, xLLM has probably trained a 375B-A23B
model but published no efficiency data, so benchmark it against
TorchTitan and Megatron-Core. For data, Datakit is the more complete
pipeline, and IFM's datasets are worth more than its toolkit. For RL,
start from upstream Miles or SkyRL rather than either lab's fork.</p>
<table>
<thead>
<tr>
<th>Dimension</th>
<th>IFM (xLLM, data toolkit, RL360)</th>
<th>Marin (Levanter + Grug, Datakit, MarinSkyRL)</th>
<th>Edge</th>
</tr>
</thead>
<tbody>
<tr>
<td>New architectures</td>
<td>Two model types chosen by flags; new blocks train through autograd,
but the fast path needs a hand-written backward</td>
<td>Copy a Grug template and edit plain JAX; autodiff supplies
gradients</td>
<td>Levanter</td>
</tr>
<tr>
<td>Optimizers</td>
<td>AdamW only</td>
<td>16, including Muon and MuonH (used by the hero)</td>
<td>Levanter</td>
</tr>
<tr>
<td>Attention and linear mixers</td>
<td>Full attention in every released model; Gekko hybrid unused; no
sliding window</td>
<td>Sliding-window/global hybrid in production; Gated DeltaNet and
Mamba-3 unwired; KDA on branches</td>
<td>Neither is broad</td>
</tr>
<tr>
<td>Long-context training</td>
<td>3.7B and 7B models mid-trained to 512K tokens</td>
<td>67B MoE continued at 262K on TPUs; 262K on GPUs only in tests</td>
<td>xLLM</td>
</tr>
<tr>
<td>FP8 / FP4</td>
<td>None</td>
<td>FP8 for classic dense layers; MoE FP8 tried and shelved; no FP4</td>
<td>Neither</td>
</tr>
<tr>
<td>GPU kernels</td>
<td>Set up for Hopper; no Blackwell run reported</td>
<td>Tuned for Blackwell; older Hopper attention; patched XLA plugin</td>
<td>Levanter on Blackwell; likely xLLM on Hopper</td>
</tr>
<tr>
<td>Non-NVIDIA hardware</td>
<td>None</td>
<td>TPU v4 to v6e</td>
<td>Levanter</td>
</tr>
<tr>
<td>Parallelism</td>
<td>FSDP, TP ≤ 8, CP; EP tied to TP; no PP</td>
<td>FSDP, EP64 in production, CP, cross-rack DP; PP in benchmarks</td>
<td>Levanter</td>
</tr>
<tr>
<td>Sparse MoE</td>
<td>Top-k routing, shared experts, no token dropping; EP within one
node; probably trained a 375B-A23B, efficiency unpublished</td>
<td>Eight transports, capacity controls, EP64; 535B-A23B hero at 26.68%
MFU</td>
<td>Levanter</td>
</tr>
<tr>
<td>Operations</td>
<td>Slurm or torchrun, a shared filesystem; resume only on the same
layout</td>
<td>Kubernetes and TPU pools; object-store checkpoints that restore onto
any mesh</td>
<td>Levanter</td>
</tr>
<tr>
<td>Performance</td>
<td>About 48.6% MFU, derived from logs, for a dense 7B over 21.9T
tokens</td>
<td>26.68% MFU for a 535B-A23B MoE on 704 GB200s</td>
<td>No like-for-like data</td>
</tr>
<tr>
<td>Evolution by agents</td>
<td>Small, but no CI or agent docs; no model module imports without a
CUDA build</td>
<td>Agent instructions, 40 skills and CPU-runnable tests; agents already
change the stack</td>
<td>Levanter</td>
</tr>
<tr>
<td>Code size</td>
<td>21.5K lines Python, 14.8K own C++/CUDA, plus 26.2K in xattn</td>
<td>84K lines Python, 1.3K CUDA, plus pinned external kernels</td>
<td>Mixed</td>
</tr>
<tr>
<td>Tests</td>
<td>62 test functions; kernels and data pipeline only; no CI</td>
<td>About 1,900 test functions; CPU and TPU tests on pull requests;
daily GPU and TPU canaries, but no GPU tests on pull requests</td>
<td>Levanter</td>
</tr>
<tr>
<td>Data curation</td>
<td>Toolkit labels, sorts and shuffles (3.1K lines, no tests); newer
sources' processing unpublished; about 13.5T tokens of data
released</td>
<td>Datakit covers download to tokenized store, including deduplication
and decontamination (360 tests); corpus not rehosted; code that chose
the hero's buckets off <code>main</code></td>
<td>Marin for code, IFM for released data</td>
</tr>
<tr>
<td>Data mixing</td>
<td>Weights passed to the trainer, none published; search method
described only in a talk, which reports small gains for 375K
GPU-hours</td>
<td>Proxy-model swarms over 200 topic × quality buckets; latest re-mix
1.20× over the previous mix on a ladder; weights published, search code
off <code>main</code></td>
<td>Marin</td>
</tr>
<tr>
<td>Post-training</td>
<td>SFT in xLLM; RL360 snapshot of Miles, Megatron-LM, SGLang, SMG and
Harbor with one demo recipe and no tests of its own; merge code private;
six post-trained models shipped, four with RL</td>
<td>SFT in JAX, partly on unmerged branches; MarinSkyRL (SkyRL fork,
Megatron only) with about 2,300 tests and nightly GPU runs; first
release pending</td>
<td>Marin for RL code, IFM for results</td>
</tr>
</tbody>
</table>
<p><em>Terms.</em> MFU: the share of peak hardware FLOPs spent on the
model's own arithmetic. FSDP, DP, TP, CP, EP, PP: fully sharded data,
data, tensor, context, expert and pipeline parallelism; EP64 means
64-way. Hopper and Blackwell: NVIDIA's H100/H200 (SM90) and B200/GB200
(SM100) generations. XLA: the compiler behind JAX. GQA: grouped-query
attention. KDA: Kimi Delta Attention, a linear-attention layer. SFT:
supervised fine-tuning. RL: reinforcement learning; RLVR uses verifiable
rewards, such as passing tests. A rollout is one sampled answer or agent
episode. GRPO and RLOO: RL methods that score each rollout against
others for the same prompt; DAPO adds asymmetric clipping and drops
uninformative prompts. TIS: truncated importance sampling, which
corrects for mismatch between the rollout and training engines. Router
replay: reusing the inference engine's expert choices during training.
On-policy distillation: training a model on its own samples, scored by a
teacher. pass@k: the chance that at least one of k samples is correct.
ISO and RAM: two weight-merging methods named in IFM's model cards. BPB:
bits per byte, a loss that does not depend on the tokenizer. MinHash and
LSH: hashing methods for finding near-duplicate documents. Miles, slime
and SkyRL are RL frameworks, SGLang and vLLM are inference engines, and
Harbor runs agent tasks in sandboxes.</p>
<h2 id="the-two-stacks">The two stacks</h2>
<p><strong>xLLM and K2 Horizon.</strong> IFM released six K2 Horizon
models on 2026-09-03: dense 0.9B, 3.7B, 7B and 32B models, MoVA-36B-A4B
(an MoE that also routes its attention value projections to experts) and
a 375B-A23B MoE, with contexts up to 512K tokens (<a
href="https://huggingface.co/IFM/K2-Horizon-7B">model card</a>). The
launch post calls xLLM <a
href="https://web.archive.org/web/20260913112353/https://ifm.ai/blog/k2/">"our
production-tested training infrastructure"</a>. IFM's public W&amp;B
logs show an internal version of xLLM training the three smallest
models: the runs use xLLM's metric names and also log metrics the public
code cannot produce (<a href="https://wandb.ai/llm360/K2-Horizon-7B">7B
logs</a>). The cards credit xLLM for the larger three as well, and a
375B run name mentions 256 nodes and 8-way expert parallelism, but no
public log confirms the trainer (section 7 weighs the evidence).
Reinforcement learning ran on separate harnesses, including internal
versions of <a href="https://github.com/ifm-ai/RL360">RL360</a>, a
Megatron-LM-based stack that IFM has released as a snapshot (section
14). IFM does not disclose K2 Horizon's hardware; logged memory peaks of
about 140 GiB per GPU point to H200s, which IFM used for its previous
model (<a href="https://arxiv.org/abs/2512.06201">K2-V2</a>).</p>
<p>The public history is two commits; the second adds 306 files and 64K
lines in one squash (<a
href="https://github.com/ifm-ai/xllm/commit/889db388d021c92164e7a675f6884dd40ea117c7">commit</a>).
GitHub lists one contributor, Xuezhe (Max) Ma, lead author of the MEGA,
Megalodon and <a href="https://arxiv.org/abs/2601.06463">Gecko</a>
architectures; xLLM's code descends from Ma's Megalodon repository. Two
companion repositories matter: <a
href="https://github.com/ifm-ai/xattn">xattn</a>, attention kernels
derived from FlashAttention-3 (<a
data-orig-href="https://github.com/ifm-ai/xattn/blob/0d6d73b37acc644d5559007fad370185b78ddbb5/THIRD_PARTY_NOTICES#L1-L16" href="https://github.com/ifm-ai/xattn/blob/0d6d73b37acc644d5559007fad370185b78ddbb5/THIRD_PARTY_NOTICES#L1-L16">notices</a>),
and <a href="https://github.com/ifm-ai/xbridges">xbridges</a>, which
converts checkpoints for Hugging Face and vLLM. xLLM needs PyTorch 2.11
or later and CUDA 12.8 or later, and launches under torchrun or Slurm
(<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/README.md#L39-L40" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/README.md#L39-L40">README</a>).</p>
<p><strong>Levanter.</strong> Two model styles share one trainer, data
loader and checkpointer. <em>Classic</em> Levanter builds models from
Haliax named axes, maps them to devices through configuration, and
imports and exports about 15 Hugging Face architectures (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/models/lm_model.py#L154-L195" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/models/lm_model.py#L154-L195">registry</a>).
<em>Grug</em>, added in January 2026, writes each model in plain JAX
with explicit sharding and asks developers to copy a template and edit
it (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/README.md#L1-L27" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/README.md#L1-L27">README</a>).
David Hall gave the reason: <a
href="https://discord.com/channels/1354881461060243556/1462884917292699669/1464708102887706725">"the
motivation is that agents will know jax way better than haliax"</a>.
Grug trains Marin's frontier runs: the 535B-A23B <em>hero</em> (704
GB200s, 7.96T of a planned 18T tokens by 2026-09-28) and Snowball
67B-A2B (10T tokens on a TPU v4-2048, <a
data-orig-href="https://github.com/marin-community/marin/issues/6044#issuecomment-5401358100" href="https://github.com/marin-community/marin/issues/6044#issuecomment-5401358100">#6044</a>).
Marin curates pretraining data with Datakit and runs RL with MarinSkyRL,
a PyTorch fork of SkyRL (sections 13 and 14).</p>
<h2 id="1-specifying-new-architectures">1. Specifying new
architectures</h2>
<p><strong>xLLM</strong> picks one of two architectures,
<code>transformer</code> or <code>gekko</code>, from a hard-coded
dictionary (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/build.py#L29-L33" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/build.py#L29-L33">build.py</a>).
Each setting is a single command-line flag; the parser rejects
dictionaries and has no syntax for per-layer lists (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/configuration/configuration.py#L140-L190" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/configuration/configuration.py#L140-L190">configuration.py</a>).
So layers cannot vary, except that the first N may be dense before the
MoE layers begin (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L607-L612" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L607-L612"><code>num_dense_layers</code></a>).
There is no interleaved MoE, no per-layer attention type, and no mixing
of Transformer and Gekko layers.</p>
<p>Each block exists twice: an eager <code>nn.Module</code> that trains
through autograd and serves evaluation, and an optional fused
<code>autograd.Function</code> with a hand-written backward. At equal
settings the eager path is nearly as fast (43.3% against 44.1% MFU on
Llama3-8B), but the fastest configuration, 52.1%, needs the fused path
and its fine-grained recomputation switches. The fused functions take 44
to 65 positional inputs, and their backward passes run 240 to 578 lines
(<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L120-L167" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L120-L167">dense
block</a>). No test compares the two paths, and they already disagree:
the fused path always uses plain additive residuals and FlashAttention,
whatever the configuration asks for (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/fused_blocks/transformer/transformer.py#L117" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/fused_blocks/transformer/transformer.py#L117">residual</a>,
<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/fused_blocks/mha.py#L73-L75" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/fused_blocks/mha.py#L73-L75">attention</a>).
A block that must also be exported needs separate Hugging Face and vLLM
ports in xbridges. MoVA came to about 2,100 lines across seven files in
two repositories.</p>
<p>xLLM trains with AdamW alone, in one parameter group, so weight decay
also applies to embeddings and norm gains (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/optim/optimizer.py#L24-L32" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/optim/optimizer.py#L24-L32">optimizer.py</a>).
There is no Muon, gradient accumulation, weight averaging, multi-token
prediction or logit z-loss; a router z-loss setting exists, but nothing
reads it (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L352" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L352">config.py</a>).</p>
<p><strong>Levanter.</strong> A Grug variant starts as a copy of a
275-line model file (<a
href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/base/model.py">base/model.py</a>),
and JAX differentiates it, so each block exists once. In unrolled
variants, per-layer choices are ordinary Python. The hero scans one
compiled block over its 48 layers and passes each layer's choice of
local or global attention in as data (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/moe_hero_ep/model.py#L1298-L1339" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/moe_hero_ep/model.py#L1298-L1339">model.py</a>).
A contract test traces one training step of each variant, except the
pipeline one, using abstract shapes on a CPU mesh (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/tests/test_grug_variant_contracts.py#L248-L288" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/tests/test_grug_variant_contracts.py#L248-L288">test</a>),
and CI posts each new variant's diff against its closest relative.
Sixteen optimizers are registered, including Muon, MuonH, SOAP and Kron
(<a
href="https://github.com/marin-community/marin/tree/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/optim">optim/</a>),
and the trainer supports z-loss and weight averaging.</p>
<p>Grug pays for this freedom in five ways. The hero's model file
repeats about 1,000 lines of its FSDP sibling, and one optimizer file
exists in five copies. Sharding lives in model code. The scanned hero
needs identical parameter shapes in every layer, so it does not support
dense layers before MoE layers. Grug cannot import Hugging Face weights.
And the classic Muon optimizers find matrices by looking for Haliax
<code>Linear</code> layers (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/optim/muon.py#L101-L121" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/optim/muon.py#L101-L121">muon.py</a>),
which Grug models lack; on a Grug model they would silently apply
AdamW-style updates, so Grug needs its own variant (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/optim/grugmuon.py#L4-L8" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/optim/grugmuon.py#L4-L8">grugmuon.py</a>).</p>
<p><em>Verdict: Levanter. In Grug a new block is one Python edit. In
xLLM a new block trains easily, but reaches full speed only through a
second, hand-differentiated implementation that nothing checks.</em></p>
<h2 id="2-attention-linear-mixers-and-long-context">2. Attention, linear
mixers and long context</h2>
<table>
<thead>
<tr>
<th>Mechanism</th>
<th>xLLM</th>
<th>Levanter + Grug</th>
</tr>
</thead>
<tbody>
<tr>
<td>Causal softmax attention, GQA, packed documents</td>
<td>FlashAttention 2, 3 or 4 (external, unpinned)</td>
<td>Splash (TPU); FlashAttention-4 via CuTe (GPU), skipping masked
document blocks</td>
</tr>
<tr>
<td>Sliding window</td>
<td>Absent: window hard-coded to −1 (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/fused_ops/attention/flash.py#L119-L131" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/fused_ops/attention/flash.py#L119-L131">flash.py</a>)</td>
<td>In production: a 2,048-token window on three of every four hero
layers</td>
</tr>
<tr>
<td>Attention sinks, logit soft-capping</td>
<td>Absent</td>
<td>Classic reference and Splash paths only</td>
</tr>
<tr>
<td>Multi-head latent attention (MLA)</td>
<td>Absent</td>
<td>Classic layer, unwired; rejected for the hero (<a
data-orig-href="https://github.com/marin-community/marin/issues/6522#issuecomment-4868376304" href="https://github.com/marin-community/marin/issues/6522#issuecomment-4868376304">#6522</a>)</td>
</tr>
<tr>
<td>Sliding-chunk attention plus delta-rule memory (Gekko)</td>
<td>Wired, IFM's own design, used by no released model</td>
<td>Absent</td>
</tr>
<tr>
<td>Gated DeltaNet, KDA</td>
<td>Absent</td>
<td>Gated DeltaNet in pure JAX, unwired; KDA on branches with an
H100-only kernel</td>
</tr>
<tr>
<td>Mamba-style state-space layers</td>
<td>Absent</td>
<td>Mamba-3 and SSD reference code, unwired</td>
</tr>
<tr>
<td>Published Gecko layer (complex moving average, Megalodon line)</td>
<td>Present, broken, unwired</td>
<td>Absent</td>
</tr>
<tr>
<td>Short causal convolution</td>
<td>In Gekko (Triton)</td>
<td>In the hero (Triton on GPU, XLA on TPU)</td>
</tr>
<tr>
<td>Context parallelism</td>
<td>Gathers all keys and values; Gekko passes state rank to rank; no
public example or test</td>
<td>Gathers all keys and values; tested at 262K tokens on 64 GB200s</td>
</tr>
<tr>
<td>RoPE scaling</td>
<td>Absent</td>
<td>Classic: YaRN and Llama 3 scaling; Grug: a temperature on
queries</td>
</tr>
</tbody>
</table>
<p>Every released K2 Horizon model is a plain Transformer with full
causal attention, and every example trains at 8,192 tokens without
context parallelism (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/k2-horizon-7B.sh#L23-L57" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/k2-horizon-7B.sh#L23-L57">example</a>).
For K2 Horizon, the README's <a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/README.md#L5" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/README.md#L5">"extra-long
contexts"</a> means full attention at long sequence lengths, and IFM has
used it at scale. The 7B was mid-trained on 1.1T tokens at 32K, 498B at
128K and 309B at 512K, then fine-tuned at 512K. Across its four 512K
phases it ran at medians of 632 to 991 tokens/s/GPU, 33–51% of peak by
xLLM's formula, which credits attention work that document masking skips
(<a href="https://wandb.ai/llm360/K2-Horizon-7B">7B logs</a>). The logs
do not say how each sequence was split across GPUs, and IFM has not
released the mid-training data.</p>
<p>Marin's longest-context training ran at 262K tokens on TPUs: about
67B tokens of continued pretraining for the 67B-A2B, then fine-tuning
(<a
data-orig-href="https://github.com/marin-community/marin/issues/6044#issuecomment-5401431533" href="https://github.com/marin-community/marin/issues/6044#issuecomment-5401431533">#6044</a>,
<a
href="https://github.com/marin-community/marin/issues/8954">#8954</a>).
On GPUs, Marin's 262K runs are tests at about 10% MFU (<a
href="https://github.com/marin-community/marin/pull/9119">#9119</a>).</p>
<p>Gekko is xLLM's most original component. Each layer normalizes with
decaying running statistics and applies short convolutions to queries,
keys and values. It then adds two branches: softmax attention over the
current and previous chunk, and an <em>adaptive working memory</em> that
summarizes older chunks with a delta-rule update (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/gated_delta_attention.py#L326-L451" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/gated_delta_attention.py#L326-L451">gated_delta_attention.py</a>).
It adapts Ma's published <a
href="https://arxiv.org/abs/2601.06463">Gecko</a> architecture, which
was pretrained at 7B on 2T tokens: it keeps Gecko's decaying norm, chunk
attention and memory update, drops its complex moving average, and adds
short convolutions. Under context parallelism each rank hands its memory
to the next. Gekko's parts have reference tests, but no test covers a
whole block and no released model uses it. Its memory update runs as a
Python loop over chunks (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/fused_ops/adaptive_working_memory.py#L99-L241" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/fused_ops/adaptive_working_memory.py#L99-L241">adaptive_working_memory.py</a>),
and because xLLM always sends a document mask, xattn's fast
chunk-attention kernel runs only on Hopper.</p>
<p>Three attention paths do not run. The <code>xattn</code> backend
raises <code>NotImplementedError</code> when a Transformer is built (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/causal_attention.py#L53-L54" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/causal_attention.py#L53-L54">causal_attention.py</a>).
The <code>swift</code> backend cannot unpack the segment data that
packed documents produce (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L719-L725" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L719-L725">producer</a>,
<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/causal_attention.py#L124-L125" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/causal_attention.py#L124-L125">consumer</a>)
and stops above 8,192 keys (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/csrc/softmax.cuh#L387-L391" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/csrc/softmax.cuh#L387-L391">softmax.cuh</a>).
And the module for the published Gecko layer calls functions with
outdated signatures (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moving_average_gated_attention.py#L199-L208" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moving_average_gated_attention.py#L199-L208">moving_average_gated_attention.py</a>).</p>
<p>Levanter's production design is softmax attention tuned by ablation:
sliding windows with a global layer every fourth layer that drops
positional encoding, query-key norms, per-head gates and short
convolutions. The same ablation found full multi-head attention better
than the grouped-query heads the hero kept (<a
data-orig-href="https://github.com/marin-community/marin/issues/8227#issuecomment-5310665092" href="https://github.com/marin-community/marin/issues/8227#issuecomment-5310665092">#8227</a>).
Its linear-attention work is younger. On 8×H100 ladder models with up to
291M active parameters, KDA in the local layers gave 1.13–1.26× the
baseline's compute efficiency, and a hybrid adding MLA global layers and
attention residuals gave 1.14–1.43× (<a
href="https://github.com/marin-community/marin/issues/9438">#9438</a>,
<a
href="https://github.com/marin-community/marin/issues/9451">#9451</a>).
That work lives on branches, with an H100 kernel and none for TPU; TPU
kernels for Gated DeltaNet are in draft (<a
href="https://github.com/marin-community/marin/pull/9350">#9350</a>).
One trap: the hero passes its window only through FlashAttention-4's
bounds, so if its model runs on the reference or Splash attention path,
the sliding-window layers silently become full attention (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/moe_hero_ep/model.py#L1283-L1329" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/moe_hero_ep/model.py#L1283-L1329">model.py</a>).</p>
<p>PyTorch offers an advantage xLLM leaves unused. The Flash Linear
Attention library has tuned kernels for dozens of linear-attention
variants, and xLLM imports it only for a causal convolution. Marin had
to write its own chunked KDA kernel for JAX in September (<a
data-orig-href="https://github.com/marin-community/marin/issues/9438#issuecomment-5826696333" href="https://github.com/marin-community/marin/issues/9438#issuecomment-5826696333">#9438</a>).</p>
<p><em>Verdict: neither stack offers a broad, production-grade set of
scalable mixers. Levanter runs the richer attention design in production
and tests more candidates; xLLM has one complete hybrid that no released
model uses. On long context, xLLM has done more in production, on
smaller models.</em></p>
<h2 id="3-fp8-and-fp4-training">3. FP8 and FP4 training</h2>
<p><strong>xLLM has no low-precision training path.</strong> The config
accepts only <code>bf16</code>, <code>fp16</code> or <code>fp32</code>
(<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L414" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L414">config.py</a>).
xLLM copied part of NVIDIA's Transformer Engine <a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/csrc/transformer_engine/NOTICE#L6-L9" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/csrc/transformer_engine/NOTICE#L6-L9">"for
its cuBLAS grouped matrix multiplication backend"</a>; the FP8 code in
that copy never runs, because xLLM passes no scaling factors (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/csrc/transformer_engine/common.cpp#L37-L61" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/csrc/transformer_engine/common.cpp#L37-L61">common.cpp</a>).
FP4 exists only as type definitions.</p>
<p><strong>Levanter has FP8 code that no production run uses.</strong>
Haliax provides an FP8 matmul with E4M3 forward values, E5M2 gradients
and per-tensor delayed scaling, wired only for classic dense
<code>Linear</code> layers (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/haliax/src/haliax/quantization.py#L180-L270" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/haliax/src/haliax/quantization.py#L180-L270">quantization.py</a>).
The MoE layer and all of Grug lack it. Marin measured MoE FP8 carefully,
found modest gains, and parked it:</p>
<ul>
<li>18B MoE, 8 H100s per arm: FP8 raised MFU from 15.1% to 15.9%, with
final loss within 0.004 (<a
data-orig-href="https://github.com/marin-community/marin/issues/7298#issuecomment-5010209991" href="https://github.com/marin-community/marin/issues/7298#issuecomment-5010209991">#7298</a>).</li>
<li>64 GB200s, single runs: MXFP8 was 1.31× faster than BF16 with 8-way
expert parallelism (<a
data-orig-href="https://github.com/marin-community/marin/issues/7282#issuecomment-5017713906" href="https://github.com/marin-community/marin/issues/7282#issuecomment-5017713906">#7282</a>)
but 0.75× as fast with FSDP alone (<a
data-orig-href="https://github.com/marin-community/marin/issues/7282#issuecomment-5037422197" href="https://github.com/marin-community/marin/issues/7282#issuecomment-5037422197">#7282</a>).</li>
<li>32 GB200s, 66B tokens: MXFP8 gave 7.2% more throughput, but final
eval loss was 0.056% worse and BF16 won all 32 paired evaluations (<a
data-orig-href="https://github.com/marin-community/marin/issues/7271#issuecomment-5036746436" href="https://github.com/marin-community/marin/issues/7271#issuecomment-5036746436">#7271</a>).</li>
<li>The owner closed the MoE FP8 pull requests in August: <a
data-orig-href="https://github.com/marin-community/marin/pull/6880#issuecomment-5402446303" href="https://github.com/marin-community/marin/pull/6880#issuecomment-5402446303">"Closing
since FP8 is not currently a priority."</a> A September retry on 8 H100s
was parked at a best gain of about 6% (<a
data-orig-href="https://github.com/marin-community/marin/issues/9438#issuecomment-5827536865" href="https://github.com/marin-community/marin/issues/9438#issuecomment-5827536865">#9438</a>),
though on 2026-09-18 David Hall wrote of FP8 expert matmuls, <a
data-orig-href="https://github.com/marin-community/marin/issues/6699#issuecomment-5727082751" href="https://github.com/marin-community/marin/issues/6699#issuecomment-5727082751">"still
think we should do this"</a>.</li>
</ul>
<p>On FP4, David Hall wrote: <a
data-orig-href="https://github.com/marin-community/marin/issues/7403#issuecomment-5050343888" href="https://github.com/marin-community/marin/issues/7403#issuecomment-5050343888">"I
strongly think we should stay away from nvfp4 unless we have very very
good reason to take on the risk"</a>. The hero computes in BF16 with
FP32 weights.</p>
<p><em>Verdict: neither trains in FP8 or FP4. Levanter has a working FP8
matmul and a record of measurements. xLLM would need new work: PyTorch's
FP8 libraries could cover its eager modules, but each fused block would
need hand changes.</em></p>
<h2 id="4-kernels-for-different-gpus">4. Kernels for different GPUs</h2>
<p><strong>xLLM</strong> leaves attention to FlashAttention: FA2 by
default, FA3 or FA4 if an environment variable is set before import (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/fused_ops/attention/flash.py#L7-L23" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/fused_ops/attention/flash.py#L7-L23">flash.py</a>).
Every example sets FA3, which runs only on Hopper. Its own 14.8K lines
of C++/CUDA mostly implement norms, moving-average scans and FFT
convolutions for the Megalodon and Gekko line, and contain no
Hopper-specific code; the hand-written kernels leave matrix products to
cuBLAS. On the released Transformer models they contribute only grouped
RMSNorm, plus an optional grouped GEMM (one cuBLASLt call per expert)
and a Triton kernel for MoE output. About half the registered ops are
never reached. The Dockerfile compiles for SM90 alone (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/Dockerfile#L22" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/Dockerfile#L22">Dockerfile</a>),
and xattn builds for SM80 through SM90 and says <a
data-orig-href="https://github.com/ifm-ai/xattn/blob/0d6d73b37acc644d5559007fad370185b78ddbb5/README.md#L84-L88" href="https://github.com/ifm-ai/xattn/blob/0d6d73b37acc644d5559007fad370185b78ddbb5/README.md#L84-L88">"SM100
and newer architectures are currently unsupported"</a>. On Blackwell,
xLLM would depend on FA4 alone; its attention test can select FA4, but
IFM reports no Blackwell run. It uses neither <code>torch.compile</code>
nor CUDA graphs.</p>
<p>xLLM's dependency risks are quieter than Levanter's but real:
FlashAttention and Flash Linear Attention are unpinned, the Transformer
Engine copy records no upstream version, its FSDP1 code imports private
PyTorch internals (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/distributed/fully_sharded_data_parallel.py#L10-L13" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/distributed/fully_sharded_data_parallel.py#L10-L13">fully_sharded_data_parallel.py</a>),
and xbridges' own vLLM path requires exactly vLLM 0.24.0.</p>
<p><strong>Levanter</strong> tunes for Blackwell. The hero uses upstream
FlashAttention-4 kernels for SM100 (<a
href="https://github.com/marin-community/marin/pull/9332">#9332</a>),
QuACK kernels for the expert matrix multiplies and for Muon, and Triton
kernels for token gathering and short convolutions. Its expert
all-to-all needs a patched XLA plugin built only for GB200's ARM hosts
(<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/pyproject.toml#L189-L192" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/pyproject.toml#L189-L192">pin</a>).
On Hopper, the attention forward pass is Marin's port of CUTLASS's
Ampere FlashAttention-2 example, without Hopper's asynchronous
tensor-core instructions (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/grug/attention/_fa4_cute_kernels.py#L32-L60" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/grug/attention/_fa4_cute_kernels.py#L32-L60">_fa4_cute_kernels.py</a>);
I found no measurement of it against FA3. TPUs use upstream Splash
attention and Megablox grouped matmul.</p>
<p>The price is maintenance. Marin pins fast-moving kernel packages
exactly (a FlashAttention-4 beta, QuACK 0.6.4, CUTLASS DSL 4.6.2),
rebuilds an XLA fork when JAX changes, and logged about a dozen kernel
and runtime incidents, from corruptions to hangs, between July and
September. Two examples: a deadlock inside a QuACK grouped GEMM that
intermittently hung the hero (<a
data-orig-href="https://github.com/marin-community/marin/issues/8870#issuecomment-5687716531" href="https://github.com/marin-community/marin/issues/8870#issuecomment-5687716531">#8870</a>),
and a Triton grouped-matmul bug that gave wrong rows on GB200 (69% of
rows in a test layer) and stayed open for five weeks (<a
href="https://github.com/marin-community/marin/issues/7484">#7484</a>,
<a href="https://github.com/marin-community/marin/pull/8610">#8610</a>).
A few engineers own the kernels, and the hero's GB200 path rests mainly
on one of them, working with an agent. IFM's incident history is
private, so the two records cannot be compared; its 7B run restarted 95
times.</p>
<p><em>Verdict: Levanter on Blackwell. On Hopper, xLLM probably has the
faster attention (FA3), though no one has measured the two side by side.
Neither stack supports AMD GPUs.</em></p>
<h2 id="5-hardware-beyond-nvidia-gpus">5. Hardware beyond NVIDIA
GPUs</h2>
<p><strong>xLLM</strong> is CUDA-only. It hard-codes NCCL, NVML
monitoring and <code>.cuda()</code> calls, loads a compiled CUDA
extension at import, and embeds NVIDIA PTX instructions in its Triton
kernels. There is no ROCm, TPU, Intel or Apple path.</p>
<p><strong>Levanter</strong> runs on TPU v4, v5e, v5p and v6e. TPU tests
run on v5e for pull requests that touch the libraries (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/.github/workflows/unified-unit.yaml#L231-L294" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/.github/workflows/unified-unit.yaml#L231-L294">workflow</a>),
a v6e canary trains daily (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/.github/workflows/marin-canary-ferry.yaml#L3-L5" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/.github/workflows/marin-canary-ferry.yaml#L3-L5">canary</a>),
and Marin trained its 8B, 32B and Snowball 67B-A2B models on TPUs. TPUs
are now secondary for Marin: Will Held wrote on 2026-09-26, <a
href="https://discord.com/channels/1354881461060243556/1553196386843893771/1553196662199947285">"We
are mostly no longer on TPUs!"</a>. The hero's expert-parallel
transport, FA4 and pipeline parallelism have no TPU versions, Marin has
no TPU v7 plan, and on AMD, <a
href="https://github.com/marin-community/marin/issues/9462">"Nobody has
run Levanter or Grug on AMD GPUs."</a></p>
<p><em>Verdict: Levanter, the only one of the two that runs anywhere but
NVIDIA.</em></p>
<h2 id="6-parallelism">6. Parallelism</h2>
<table>
<thead>
<tr>
<th></th>
<th>xLLM</th>
<th>Levanter + Grug</th>
</tr>
</thead>
<tbody>
<tr>
<td>Data parallel / FSDP</td>
<td>FSDP1 or FSDP2 (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/build.py#L73-L150" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/build.py#L73-L150">build.py</a>)</td>
<td>FSDP within a rack or slice; plain data parallel across racks or TPU
slices</td>
</tr>
<tr>
<td>Tensor parallel</td>
<td>Megatron-style; up to 8 ways in practice</td>
<td>Available in classic; unused in the hero by choice (<a
data-orig-href="https://github.com/marin-community/marin/issues/6367#issuecomment-4888720086" href="https://github.com/marin-community/marin/issues/6367#issuecomment-4888720086">#6367</a>)</td>
</tr>
<tr>
<td>Context parallel</td>
<td>Gathers all keys and values; no public example or test</td>
<td>Gathers all keys and values; 262K-token tests on GPU (<a
href="https://github.com/marin-community/marin/pull/9119">#9119</a>);
production on TPU</td>
</tr>
<tr>
<td>Expert parallel</td>
<td>Uses the tensor-parallel group, so at most 8 ranks within one
node</td>
<td>Its own mesh axis; 64 ranks across a GB200 rack in production</td>
</tr>
<tr>
<td>Pipeline parallel</td>
<td>Absent</td>
<td>JaxPP benchmarks only (<a
href="https://github.com/marin-community/marin/pull/8739">#8739</a>, <a
href="https://github.com/marin-community/marin/issues/9277">#9277</a>)</td>
</tr>
<tr>
<td>Gradient accumulation</td>
<td>Absent</td>
<td>Classic yes; Grug no</td>
</tr>
<tr>
<td>Largest run</td>
<td>About 600 GPUs for the 7B, inferred from logs; about 2,048 for the
375B, inferred from a run name</td>
<td>704 GB200s in production since August</td>
</tr>
</tbody>
</table>
<p>In xLLM, each tensor-parallel rank holds a slice of every token's
hidden vector. An all-to-all swaps those slices so that each rank sees
whole vectors for its own experts (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/moe.py#L165-L195" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/moe.py#L165-L195">moe.py</a>).
Tokens never move between data-parallel ranks, so expert parallelism
cannot exceed the tensor-parallel degree, in practice eight GPUs. On
8-GPU H200 nodes that is a common layout, and the 375B run name says
<code>ep8</code>. But FSDP must then gather each rank's full set of
local experts at every layer, and the design cannot use a 72-GPU NVLink
domain. Marin measured this choice on one GB200 rack: at matched token
drops, expert parallelism reached 22.90% MFU against 19.40% for FSDP at
the same shape (<a
href="https://github.com/marin-community/marin/pull/7981">#7981</a>).</p>
<p>Levanter's pipeline and context parallelism work in benchmarks but
not yet in GPU production; drafts extend both (<a
href="https://github.com/marin-community/marin/pull/9279">#9279</a>, <a
href="https://github.com/marin-community/marin/pull/9460">#9460</a>, <a
href="https://github.com/marin-community/marin/pull/9282">#9282</a>).
The Grug hero has no gradient accumulation, so its batch size is tied to
the number of racks, and its expert parallelism stops at the rack
boundary. On H100s, Marin also uses 8-way expert parallelism within a
node, as xLLM does.</p>
<p><em>Verdict: Levanter for large MoE models. For dense models on 8-GPU
nodes, xLLM's FSDP-first design has proven sufficient.</em></p>
<h2 id="7-mixture-of-experts-support-and-the-375b-model">7.
Mixture-of-experts support and the 375B model</h2>
<p><strong>xLLM supports sparse MoE.</strong> It ships presets for
Mixtral-8x7B, K2 Horizon MoE-36B-A4B, MoVA-36B-A4B and the 375B-A23B (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L877-L898" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L877-L898">presets</a>).
A router picks the top k experts from sigmoid or softmax scores. A
per-expert bias, nudged toward balanced loads at every step, shifts only
which experts are chosen, and an optional <code>dot</code> or
<code>entropy</code> balancing loss can be added (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/router.py#L159-L195" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/router.py#L159-L195">router.py</a>).
Shared experts and leading dense layers are supported, and routing drops
no tokens. Experts run through one of four grouped-GEMM backends,
including a Transformer Engine grouped GEMM spread over several CUDA
streams (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/fused_ops/mgmm/te_mgmm.py#L23-L60" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/fused_ops/mgmm/te_mgmm.py#L23-L60">te_mgmm.py</a>),
and a fused MoE block with a hand-written backward is optional. MoVA
also routes the attention value projections to experts.</p>
<p>The limits matter at frontier scale:</p>
<ul>
<li>Expert parallelism uses the tensor-parallel group, so it spans at
most 8 GPUs in one node (section 6).</li>
<li>Every MoE layer stops to copy expert counts to the host, in both the
eager and fused paths (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/moe.py#L177-L178" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/moe.py#L177-L178">moe.py</a>,
<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/fused_blocks/moe.py#L90-L91" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/fused_blocks/moe.py#L90-L91">fused
block</a>).</li>
<li>There are no capacity controls and no group-limited routing, and the
router z-loss setting is never read.</li>
<li>Tests cover only the grouped-GEMM and permutation kernels. The
router test checks a copy of the router, and nothing tests a whole MoE
layer or expert-parallel correctness (<a
href="https://github.com/ifm-ai/xllm/tree/889db388d021c92164e7a675f6884dd40ea117c7/tests/moe">tests/moe</a>).</li>
<li>Gekko layers have no MoE variant (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/gekko.py#L345-L348" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/gekko.py#L345-L348">gekko.py</a>).</li>
<li>The only published MoE throughput is MoVA-36B-A4B at a claimed 26.7%
MFU on 128 H200s.</li>
</ul>
<p>Grug, for comparison, offers eight MoE backends (five of them
expert-parallel), capacity accounting, balancing by per-expert quantile
thresholds, and the 64-way expert parallelism across a GB200 rack that
the hero runs (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/grug/_moe/common.py#L87-L111" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/grug/_moe/common.py#L87-L111">common.py</a>,
<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/grug/grug_moe.py#L188-L344" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/grug/grug_moe.py#L188-L344">grug_moe.py</a>).</p>
<p><strong>Did xLLM train the 375B?</strong> Probably an internal
version of it, but the public record does not prove it.</p>
<p>Evidence for:</p>
<ul>
<li>IFM says so. The <a
href="https://huggingface.co/IFM/K2-Horizon-375B-A23B/blob/82af3bc5fbdbc24f1d15a081ac035396f9b2c3dd/README.md">model
card</a> lists <code>ifm-ai/xllm</code> as its code repository and
advises: <a
href="https://huggingface.co/IFM/K2-Horizon-375B-A23B/blob/82af3bc5fbdbc24f1d15a081ac035396f9b2c3dd/README.md">"To
continue the original pretraining with xLLM's training behavior, use the
native xLLM checkpoint and XLLM runtime."</a></li>
<li>The public preset matches the released <a
href="https://huggingface.co/IFM/K2-Horizon-375B-A23B/blob/82af3bc5fbdbc24f1d15a081ac035396f9b2c3dd/config.json">config</a>
in layer count, leading dense layers, width, attention heads, expert
count and size, sigmoid router with bias and 2.5 scaling, and norm
groups. Only the RoPE base differs, 500K in the preset against 10M in
the release, which fits a change during the move to 512K context.</li>
<li>The run ID for Pretraining Phase 2,
<code>k2moe375B_txt360v2.3_256nodes_seed42_bsz32M_seq8k_jais250k_ep8_dot_te_phase2</code>,
uses xLLM's own option values. <code>dot</code> is a value of
<code>moe_router_load_balancing_type</code> and <code>te</code> of
<code>moe_expert_backend</code> (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L421-L422" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L421-L422">config.py</a>,
<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L197-L198" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L197-L198">config.py</a>);
Megatron-LM's balancing options include no <code>dot</code> (<a
data-orig-href="https://github.com/NVIDIA/Megatron-LM/blob/73813c854b70afa0ba77ebc76fae8950213ecb20/megatron/training/arguments.py#L3748-L3749" href="https://github.com/NVIDIA/Megatron-LM/blob/73813c854b70afa0ba77ebc76fae8950213ecb20/megatron/training/arguments.py#L3748-L3749">arguments.py</a>).
<code>ep8</code> fits xLLM's tensor-parallel expert layout, which the
preset permits (8 key/value heads, one norm group), and 256 nodes pass
xLLM's startup check.</li>
<li>The 375B's rebuilt <a
href="https://wandb.ai/llm360/K2-Horizon-375B">W&amp;B curves</a> use
xLLM's metric names, <code>optim/avg_loss</code>,
<code>optim/g_norm</code> and <code>optim/lr</code> (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/train.py#L480-L499" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/train.py#L480-L499">train.py</a>).
The Phase 1 curves came from a folder named
<code>phase1_xllm_eval_tag_inventory/tb</code>, the kind of
<code>tb</code> directory xLLM writes (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/metrics.py#L40" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/metrics.py#L40">metrics.py</a>).</li>
<li>The released model code says it returns <a
href="https://huggingface.co/IFM/K2-Horizon-375B-A23B/blob/82af3bc5fbdbc24f1d15a081ac035396f9b2c3dd/modeling_k2_horizon.py">"native-XLLM-compatible
routing weights"</a>, and the config follows the schema of xbridges'
xLLM-to-Hugging Face converter.</li>
</ul>
<p>Gaps:</p>
<ul>
<li>The 375B's W&amp;B project holds no training telemetry. Its eight
runs were rebuilt after the fact from an evaluation spreadsheet and a
TensorBoard export, and record only loss, gradient norm, learning rate
and evaluations. Unlike the 0.9B, 3.7B and 7B logs, they carry no
throughput, token counter or configuration, so the run's speed and
hardware are unknown; <code>256nodes</code> in the run name implies
about 2,048 GPUs.</li>
<li>No native xLLM checkpoint, 375B launch script or technical report is
public.</li>
<li>The released config does not exactly match the public converter's
output: it lacks the converter's <code>xllm_model_parallel_size</code>
field and records bfloat16 where the converter writes float32 (<a
data-orig-href="https://github.com/ifm-ai/xbridges/blob/227f3fd26e30b024ce09bc7c556c64616c13276e/xbridges/huggingface/xllm_to_hf_main.py#L185-L187" href="https://github.com/ifm-ai/xbridges/blob/227f3fd26e30b024ce09bc7c556c64616c13276e/xbridges/huggingface/xllm_to_hf_main.py#L185-L187">xllm_to_hf_main.py</a>).</li>
<li>IFM had Megatron at hand: it pretrained its previous flagship, K2-V2
70B, with Megatron-Core (<a
href="https://arxiv.org/abs/2512.06201">K2-V2</a>), and it runs
reinforcement learning on Megatron-LM.</li>
</ul>
<p><em>Verdict: xLLM supports MoE, and an internal version of it
probably pretrained the 375B-A23B on about 2,048 GPUs for 15T tokens.
Its MoE design suits 8-GPU nodes but cannot use a larger NVLink domain,
and IFM has not published how efficiently it trained at that size.
Levanter's hero remains the only published MoE throughput at frontier
scale.</em></p>
<h2 id="8-operations">8. Operations</h2>
<ul>
<li><strong>Scheduler and storage.</strong> xLLM launches under torchrun
or Slurm, and its example scripts, requeue handler and asynchronous
evaluation assume Slurm (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/distributed/slurm.py#L16-L48" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/distributed/slurm.py#L16-L48">slurm.py</a>).
It reads and writes a shared POSIX filesystem and has no object-store
support. Marin runs Levanter through its own Iris scheduler on
Kubernetes and TPU pools, and writes checkpoints with TensorStore to
object storage (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/tensorstore_serialization.py#L942-L1040" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/tensorstore_serialization.py#L942-L1040">tensorstore_serialization.py</a>).
Levanter also detects Slurm jobs (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/distributed.py#L27-L31" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/distributed.py#L27-L31">distributed.py</a>).</li>
<li><strong>Checkpoint and resume.</strong> xLLM saves optimizer state
per rank by default, so a resume must keep the same layout: a new
tensor-parallel degree raises <code>NotImplementedError</code> (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/reloading.py#L150-L162" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/reloading.py#L150-L162">reloading.py</a>),
and a new data-parallel size fails in the data iterator (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/data/dataset_streamer/data_iterator/data_iterator.py#L189-L191" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/data/dataset_streamer/data_iterator/data_iterator.py#L189-L191">data_iterator.py</a>).
Turning on its asynchronous checkpointer fails at startup, because the
class leaves abstract methods unimplemented (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/checkpointing.py#L105-L132" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/checkpointing.py#L105-L132">checkpointing.py</a>).
Levanter writes asynchronously and restores any checkpoint onto any
mesh.</li>
<li><strong>Fault tolerance.</strong> In the public release, Slurm's
warning signal makes xLLM requeue the job without saving, and the only
hang detection is PyTorch's NCCL watchdog. IFM's production runs had
more. Their logs record GPU hardware-error checks and NCCL benchmarks
that the public code cannot produce, and IFM's K2-V2 paper describes an
in-house tool that detects hardware faults and recovers automatically
(<a href="https://arxiv.org/abs/2512.06201">K2-V2</a>). The public code
also differs in small ways from what ran: its three 38-node example
scripts enable a startup bandwidth check that requires a multiple of
four nodes, so as published they stop before training (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/cluster_check/comms_bench.py#L35-L58" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/cluster_check/comms_bench.py#L35-L58">check</a>,
<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/k2-horizon-7B.sh#L7" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/k2-horizon-7B.sh#L7">script</a>).
Levanter adds a step watchdog (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/callbacks/progress_watchdog.py#L33" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/callbacks/progress_watchdog.py#L33">progress_watchdog.py</a>),
an endpoint that forces a checkpoint, hourly checkpoints on the hero and
automatic retries, so a failure costs the hero at most about an hour of
work. Neither saves a checkpoint when warned of preemption.</li>
<li><strong>Data.</strong> xLLM tokenizes JSONL during training and
packs documents by best fit, so it needs no preprocessing step. Levanter
reads pre-tokenized caches with staged mixtures.</li>
<li><strong>MoE routing.</strong> xLLM drops no tokens, but synchronizes
with the host at every MoE layer. Grug drops assignments above a
capacity factor of 1.15: about 0.01% on the hero at 4K tokens, though in
earlier tests drops climbed from about 7% to 40% when sequences grew
from 4K to 65K tokens (<a
href="https://github.com/marin-community/marin/issues/8435">#8435</a>).
The hero trains with a logit z-loss; xLLM has none.</li>
<li><strong>Determinism.</strong> Levanter claims bitwise determinism on
TPU, even across preemption and resume (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/README.md#L43" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/README.md#L43">README</a>).
Neither stack promises it on GPUs. xLLM's <code>deterministic</code>
flag is off by default and does not reach its Triton MoE kernel, which
adds router gradients with atomic operations (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/permute/permute_kernels.py#L103-L108" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/permute/permute_kernels.py#L103-L108">permute_kernels.py</a>).</li>
</ul>
<p><em>Verdict: Levanter.</em></p>
<h2 id="9-performance">9. Performance</h2>
<table>
<thead>
<tr>
<th>Stack</th>
<th>Model</th>
<th>Hardware</th>
<th>Seq.</th>
<th>MFU</th>
<th>Status</th>
</tr>
</thead>
<tbody>
<tr>
<td>xLLM</td>
<td>K2 Horizon 7B</td>
<td>About 600 GPUs, likely H200 (inferred)</td>
<td>8K</td>
<td>≈48.6% (median 8,735 tokens/s/GPU)</td>
<td>Production, 21.9T tokens (<a
href="https://wandb.ai/llm360/K2-Horizon-7B">logs</a>)</td>
</tr>
<tr>
<td>xLLM</td>
<td>K2 Horizon 7B, four long-context phases</td>
<td>Same</td>
<td>512K</td>
<td>33–51% (medians 632–991 tokens/s/GPU)</td>
<td>Production, 558B tokens (same logs)</td>
</tr>
<tr>
<td>xLLM</td>
<td>Llama3-8B</td>
<td>128 H200, FSDP2</td>
<td>8K</td>
<td>43.3% eager; 44.1% fused; 52.1% fused at 109 GiB/GPU</td>
<td>Claimed (<a
href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/benchmark_llama3_h200_20260925.md">file</a>)</td>
</tr>
<tr>
<td>xLLM</td>
<td>MoVA-36B-A4B</td>
<td>128 H200, FSDP 64 × TP 2</td>
<td>8K</td>
<td>26.7%</td>
<td>Claimed (<a
href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/benchmark_k2-horizon_h200_20260925.md">file</a>)</td>
</tr>
<tr>
<td>Grug</td>
<td>535B-A23B hero</td>
<td>704 GB200, EP64 × DP11</td>
<td>4K</td>
<td>26.68% (median)</td>
<td>Production (<a
data-orig-href="https://github.com/marin-community/marin/issues/8317#issuecomment-5821773936" href="https://github.com/marin-community/marin/issues/8317#issuecomment-5821773936">#8317</a>)</td>
</tr>
<tr>
<td>Grug</td>
<td>Snowball 67B-A2B</td>
<td>TPU v4-2048</td>
<td>8K</td>
<td>18.6% (full-attention count)</td>
<td>Production, 10T tokens (<a
data-orig-href="https://github.com/marin-community/marin/issues/6044#issuecomment-4812848295" href="https://github.com/marin-community/marin/issues/6044#issuecomment-4812848295">#6044</a>)</td>
</tr>
<tr>
<td>Grug</td>
<td>About 45B MoE</td>
<td>64 H100</td>
<td>—</td>
<td>24.3%</td>
<td>Benchmark (<a
data-orig-href="https://github.com/marin-community/marin/issues/6979#issuecomment-4972374164" href="https://github.com/marin-community/marin/issues/6979#issuecomment-4972374164">#6979</a>)</td>
</tr>
<tr>
<td>Grug</td>
<td>Snowball 67B-A2B</td>
<td>64 H100, PP8 × EP8</td>
<td>8K</td>
<td>19.5% (full-attention count)</td>
<td>14 synthetic steps (<a
href="https://github.com/marin-community/marin/pull/8739">#8739</a>)</td>
</tr>
<tr>
<td>Levanter</td>
<td>Llama-3.1-8B</td>
<td>TPU v5p-8</td>
<td>1K</td>
<td>48.6–50.3%</td>
<td>Microbenchmark (<a
data-orig-href="https://github.com/marin-community/marin/issues/1864#issuecomment-3508681402" href="https://github.com/marin-community/marin/issues/1864#issuecomment-3508681402">#1864</a>)</td>
</tr>
<tr>
<td>Grug</td>
<td>2.7B dense</td>
<td>8 H100</td>
<td>2K</td>
<td>56.2%</td>
<td>29 steps, on a branch (<a
data-orig-href="https://github.com/marin-community/marin/issues/6979#issuecomment-4894291032" href="https://github.com/marin-community/marin/issues/6979#issuecomment-4894291032">#6979</a>)</td>
</tr>
</tbody>
</table>
<p>IFM's production medians sit within 94–101% of its benchmark claims
for the same models, so the benchmark file looks representative. Three
cautions apply.</p>
<ol type="1">
<li><strong>The stacks count FLOPs differently.</strong> xLLM halves
attention for causal masking, leaves out embedding parameters (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L788-L833" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L788-L833">formula</a>),
and credits attention that document masking skips. Levanter's generic
formula counts attention in full (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/utils/flop_utils.py#L39-L52" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/utils/flop_utils.py#L39-L52">flop_utils.py</a>):
on a Llama3-8B shape it gives the same run 1.13× xLLM's MFU at 8K and
1.8× at 262K. The hero's counter accounts for sliding windows and sits
within about 2% of a causal count at 4K. The two Snowball rows price
every layer as full attention, though most use a 2K window; counted the
hero's way, they would read about 15–16%.</li>
<li><strong>xLLM's TorchTitan comparison mixes hardware.</strong> xLLM's
file says its TorchTitan row (6,514 tokens/s/GPU) comes from
TorchTitan's documentation, but not that TorchTitan measured it in
December 2024 on 128 H100s capped at 500 W, at half the batch size (<a
data-orig-href="https://github.com/pytorch/torchtitan/blob/d43271c3303d6aa311d74ba27fa116be77b89973/benchmarks/llama3_h100_202412_torchtitan.md#L40-L46" href="https://github.com/pytorch/torchtitan/blob/d43271c3303d6aa311d74ba27fa116be77b89973/benchmarks/llama3_h100_202412_torchtitan.md#L40-L46">TorchTitan</a>);
the file's setup says <a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/benchmark_llama3_h200_20260925.md#L7" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/benchmark_llama3_h200_20260925.md#L7">"All
jobs are running 128 H200 GPUs"</a>.</li>
<li><strong>The 52.1% row needs H200 memory.</strong> It uses 109 GiB
per GPU, and the published script turns every recomputation switch off
(<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/llama3-8B.sh#L52-L68" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/llama3-8B.sh#L52-L68">script</a>).
Despite its <code>+ recompute</code> label, the row most likely ran
without recomputation.</li>
</ol>
<p><strong>External baselines.</strong> Marin has calibrated Grug
against Megatron. On 8 B200s, Grug reached 90% of Megatron's throughput
(<a
data-orig-href="https://github.com/marin-community/marin/issues/6139#issuecomment-4619246391" href="https://github.com/marin-community/marin/issues/6139#issuecomment-4619246391">#6139</a>).
On one GH200, Megatron was about 28% faster on a small MoE (<a
data-orig-href="https://github.com/marin-community/marin/issues/4311#issuecomment-4414429663" href="https://github.com/marin-community/marin/issues/4311#issuecomment-4414429663">#4311</a>).
On a GB200 rack, Megatron-Core reached 31–32% MFU with artificially
balanced routing, 26.6% with Grug's operators swapped in (still
balanced), and 14.7% with its stock learned router (<a
data-orig-href="https://github.com/marin-community/marin/issues/7668#issuecomment-5104907558" href="https://github.com/marin-community/marin/issues/7668#issuecomment-5104907558">#7668</a>).
The hero reaches 26.7% across 11 racks with trained routing, though at a
different shape.</p>
<p><em>Verdict: no like-for-like data exists. xLLM's dense numbers are
strong and hold up in production; Marin has no comparable dense
measurement. Levanter's MoE figure is the only published one at frontier
scale; IFM has published none for its 375B. One run would settle the
dense question: Llama3-8B at 8K on the same 64–128 H100s or H200s in
both stacks, scored with one formula.</em></p>
<h2 id="10-evolution-by-agents">10. Evolution by agents</h2>
<p><strong>xLLM</strong> is small and written in PyTorch, which coding
agents know best. But it gives an agent little to steer by: no
AGENTS.md, no CI, and only 6 of its 36 test files run cleanly under
pytest. No model module imports without the compiled extension,
FlashAttention and Flash Linear Attention, so an agent without a CUDA
build cannot run a single model test. The paired eager and fused
implementations, with hand-written backward passes and no parity test,
make each architecture change risky.</p>
<p><strong>Levanter</strong> was reshaped for agents. The Marin monorepo
layers instructions in AGENTS.md files and carries 40 agent skills (for
example <code>change-grug</code>, <code>add-pallas-kernel</code> and
<code>deploy-hero-change</code>); Levanter adds contract tests that
trace every variant on CPU, multi-device tests on simulated CPU devices,
and agent-run lint and review. In September, agents wrote 79% of new
issues, and agent accounts made 28% of training-stack commits; people
also run agents under their own names, so the true share is higher. An
agent decoded GPU hang dumps to trace the QuACK deadlock (<a
data-orig-href="https://github.com/marin-community/marin/issues/8870#issuecomment-5687716531" href="https://github.com/marin-community/marin/issues/8870#issuecomment-5687716531">#8870</a>)
and ran the hero's code handoffs (<a
data-orig-href="https://github.com/marin-community/marin/issues/8506#issuecomment-5804146010" href="https://github.com/marin-community/marin/issues/8506#issuecomment-5804146010">#8506</a>).
Another runs small-model experiments that anyone can steer by commenting
(<a
href="https://github.com/marin-community/marin/issues/9451">#9451</a>).
David Hall noted that <a
href="https://discord.com/channels/1354881461060243556/1462884917292699669/1464707410533941510">"grug
itself was mostly done by codex"</a>.</p>
<p>Agents still cannot validate GPU kernels, the patched transport or
multi-host behavior without hardware. An agent-run hero handoff filled
the 100 TiB storage quota and cost about two hours and 377 retrained
steps (<a
data-orig-href="https://github.com/marin-community/marin/issues/8506#issuecomment-5817400840" href="https://github.com/marin-community/marin/issues/8506#issuecomment-5817400840">#8506</a>).
Reviewers complain that agents <a
href="https://github.com/marin-community/marin/issues/8738">"produce
'slop' in a way and scope well beyond a normal human coder"</a>.</p>
<p><em>Verdict: Levanter, by a wide margin.</em></p>
<h2 id="11-code-size-and-complexity">11. Code size and complexity</h2>
<p>Counts cover whole libraries, including data, evaluation and export
code, and exclude blank and comment lines.</p>
<table>
<thead>
<tr>
<th></th>
<th>xLLM</th>
<th>Levanter + Haliax + Grug</th>
</tr>
</thead>
<tbody>
<tr>
<td>Python</td>
<td>21.5K (including 2.2K Triton)</td>
<td>57.4K + 9.5K + 17.3K</td>
</tr>
<tr>
<td>Own C++/CUDA</td>
<td>14.8K, plus 26.2K in xattn (11.5K of it modified FlashAttention
code)</td>
<td>1.3K (a DeepEP binding)</td>
</tr>
<tr>
<td>Vendored or pinned kernel code</td>
<td>3.4K lines of Transformer Engine</td>
<td>External packages: FA4, QuACK, CUTLASS DSL, patched XLA plugin</td>
</tr>
<tr>
<td>Mean cyclomatic complexity (Python)</td>
<td>3.6</td>
<td>3.0 / 2.6 / 3.2</td>
</tr>
<tr>
<td>Functions with complexity above 20</td>
<td>2.0%</td>
<td>0.7% / 0.4% / 1.4%</td>
</tr>
<tr>
<td>Largest training-loop function</td>
<td><code>main</code>: complexity 63, 461 lines (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/train.py#L200-L660" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/train.py#L200-L660">train.py</a>)</td>
<td>Hero <code>_run_grug_local</code>: complexity 65, about 390 lines
(<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/moe_hero_ep/train.py#L933-L1321" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/moe_hero_ep/train.py#L933-L1321">train.py</a>)</td>
</tr>
</tbody>
</table>
<p>Only Gekko uses xattn. xLLM's complexity sits in its twin
implementations, hand-written backward passes and dead code: about half
its CUDA ops are unreachable, and the published Gecko layer's module is
broken. Levanter's is spread across two model styles, nine attention
backends, eight MoE transports, copied variants and external runtime
pins.</p>
<p><em>Verdict: mixed. xLLM is smaller; Levanter's functions are
simpler, but it has more parts, many outside the repository.</em></p>
<h2 id="12-test-coverage">12. Test coverage</h2>
<p><strong>xLLM</strong> has 62 test functions. About 17 files compare a
kernel's forward and backward passes against a reference, 10 only time
kernels, and three CPU files (23 functions) cover document packing, data
validation and data-loader resume. Nothing tests FSDP, tensor, context
or expert parallelism, checkpoint reloading, MoE layers, the optimizer
or end-to-end loss. The router test checks a local copy of the router
that has already drifted from the production code (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/tests/moe/test_router.py#L1-L7" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/tests/moe/test_router.py#L1-L7">test</a>,
<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/router.py#L98-L99" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/router.py#L98-L99">production</a>),
and the resume smoke test passes flags that no longer exist. There is no
CI.</p>
<p><strong>Levanter</strong> has about 1,300 test functions and Haliax
340, and Marin's root suite adds about 230 in files that exercise Grug
or Snowball. They include Hugging Face parity tests,
kernel-versus-reference tests and multi-device tests on simulated CPU
devices. Pull requests run CPU shards and a TPU lane; canaries train a
Grug MoE daily on TPU v6e, 8×H100 and TPU multislice. The gap is GPUs on
pull requests: <a
data-orig-href="https://github.com/marin-community/marin/issues/8704#issuecomment-5443649625" href="https://github.com/marin-community/marin/issues/8704#issuecomment-5443649625">"no
pytest marker anywhere selects GPU work"</a>, the TPU lane ignores
changes under <code>experiments/grug</code>, and the hero's transport is
validated by hand.</p>
<p><em>Verdict: Levanter.</em></p>
<h2 id="13-data-curation-and-mixing">13. Data curation and mixing</h2>
<p>IFM's <a
href="https://github.com/ifm-ai/pretraining-data-toolkit/tree/cd93e116a8217c9e9a45e104b0ea28ac653a998f">pretraining-data-toolkit</a>
and Marin's <a
href="https://github.com/marin-community/marin/tree/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/datakit">Datakit</a>
both sort pretraining documents by topic and quality. They differ in how
much of the pipeline they cover.</p>
<table>
<thead>
<tr>
<th>Stage</th>
<th>IFM toolkit</th>
<th>Marin Datakit</th>
</tr>
</thead>
<tbody>
<tr>
<td>Sources</td>
<td>Documents prepared upstream, such as TxT360's Common Crawl text;
adapters for HPLT, S2ORC, MegaMath, PhilPapers and USPTO</td>
<td>152 open datasets in 292 registry entries, nearly all pinned Hugging
Face revisions (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/src/marin/datakit/sources.py#L141-L635" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/src/marin/datakit/sources.py#L141-L635">sources.py</a>)</td>
</tr>
<tr>
<td>Language ID, rule-based filters, PII removal</td>
<td>Absent; TxT360's public pipeline has language ID and filters for
Common Crawl</td>
<td>Absent as stages; language comes from upstream metadata</td>
</tr>
<tr>
<td>Deduplication</td>
<td>Absent; reads TxT360's duplicate counts, then drops them</td>
<td>Exact by content hash, then MinHash with LSH over character 5-grams
(<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/src/marin/processing/classification/deduplication/fuzzy_minhash.py#L145-L172" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/src/marin/processing/classification/deduplication/fuzzy_minhash.py#L145-L172">fuzzy_minhash.py</a>)
and a containment check before removal (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/src/marin/processing/classification/deduplication/cluster_dedup.py#L33-L53" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/src/marin/processing/classification/deduplication/cluster_dedup.py#L33-L53">cluster_dedup.py</a>);
fuzzy removal skips 16 mostly synthetic sources</td>
</tr>
<tr>
<td>Decontamination</td>
<td>Absent</td>
<td>13-gram Bloom filter against the nine Artificial Analysis
Intelligence Index benchmarks and lm-eval tasks (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/src/marin/datakit/decon.py#L102-L130" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/src/marin/datakit/decon.py#L102-L130">decon.py</a>)</td>
</tr>
<tr>
<td>Quality</td>
<td>Three public classifiers (FineWeb-Edu, DCLM, PreSelect); each score
is cut into 20 levels, and the highest sets one of five categories</td>
<td>A small classifier trained on LLM labels; five bands, used for
mixing rather than filtering</td>
</tr>
<tr>
<td>Topic</td>
<td>WebOrganizer's public topic and format labels, kept as columns</td>
<td>Embeddings from Microsoft's Harrier model, clustered into 5,000
groups and merged into 40 topics</td>
</tr>
<tr>
<td>Output</td>
<td>Shuffled JSONL for each quality category</td>
<td>200 tokenized Levanter caches, one per topic and band</td>
</tr>
<tr>
<td>Mixing</td>
<td>Weights passed to xLLM at launch (<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L319" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L319">config.py</a>);
none published</td>
<td>Staged mixtures in Levanter; weights from proxy-model swarms,
published; search code off <code>main</code></td>
</tr>
<tr>
<td>Orchestration</td>
<td>Scripts run by hand, locally or as Slurm arrays; labeling needs
GPUs</td>
<td>Steps keyed by a hash of their inputs, run on Marin's Zephyr
library, which needs Marin's Iris scheduler beyond one machine; daily
and weekly canaries</td>
</tr>
<tr>
<td>Tests</td>
<td>None; no CI</td>
<td>360 test functions in <code>tests/datakit</code>, plus the engine's
and Levanter's data tests</td>
</tr>
<tr>
<td>Size</td>
<td>About 3,100 lines of Python</td>
<td>About 28,900 lines; 44,600 with the engine</td>
</tr>
<tr>
<td>Data released</td>
<td>About 13.5T tokens in five datasets (my rough estimate)</td>
<td>Corpus not rehosted; recipes download each source from its
creator</td>
</tr>
</tbody>
</table>
<p><strong>IFM's toolkit is the middle of a pipeline.</strong> It labels
documents with five public classifiers, counts tokens with the JAIS-13b
tokenizer, sorts documents into quality categories, shuffles them and
inspects the result. The sorting step reuses the recipe that NVIDIA's
Nemotron-CC dataset used to combine classifier scores, including its
category names and boundaries (<a
data-orig-href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/annotate/classify.py#L176-L234" href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/annotate/classify.py#L176-L234">classify.py</a>).
Its 16 output columns match the 770-million-row
<code>web-high-medium</code> subset of <a
href="https://huggingface.co/datasets/IFM/TxT360-v2">TxT360-v2</a> in
name and order, so this code most likely produced that subset. The
public version was repackaged without being run: <a
data-orig-href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/README.md#L7-L10" href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/README.md#L7-L10">"This
environment could not install the data-processing dependencies or run
CUDA inference"</a>. An agent ran its CPU stages on synthetic data: the
shuffle kept every row and reran identically, and the category shares
came out as expected. Reading the code turns up four problems:</p>
<ul>
<li>WebOrganizer inputs are padded to 8,192 tokens by default, so by my
estimate a 1,000-token page costs about 16 times the compute it needs
(<a
data-orig-href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/annotate/annotate_v2.py#L1060-L1075" href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/annotate/annotate_v2.py#L1060-L1075">annotate_v2.py</a>).</li>
<li>After a model worker fails, the loop moves on to the next file
without clearing queued results, so later files can fail or, rarely,
receive another file's labels (<a
data-orig-href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/annotate/annotate_v2.py#L1185-L1196" href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/annotate/annotate_v2.py#L1185-L1196">annotate_v2.py</a>).</li>
<li>Sorting drops URLs, document IDs and TxT360's duplicate counts, so
released rows cannot be traced to their sources or upsampled by
duplicate count, as IFM did for K2-V2.</li>
<li>The shuffle writes <code>part-*.jsonl</code> files, but xLLM's
loader reads only file names containing <code>chunk</code> and a number
(<a
data-orig-href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/shuffle/shuffle_2ndpass.py#L44" href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/shuffle/shuffle_2ndpass.py#L44">toolkit</a>,
<a
data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/data/dataset_streamer/data_iterator/data_iterator.py#L100-L101" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/data/dataset_streamer/data_iterator/data_iterator.py#L100-L101">xLLM</a>).</li>
</ul>
<p>For K2 Horizon's newer web sources, such as ClueWeb22 and HPLT,
nothing before labeling is public: no extraction, language ID,
deduplication or decontamination. Its Common Crawl text is older. It
comes from <a
href="https://github.com/LLM360/TxT360/tree/07d98df82292651c2cf56753f9563338ab0c6073">TxT360</a>,
whose extraction, fastText language ID, filtering and global MinHash
deduplication IFM published in 2024 under its former name, LLM360, and
documented in its <a href="https://arxiv.org/abs/2512.06201">K2-V2
paper</a>. Other IFM repositories cover parts of synthesis and data
checking. <a
href="https://github.com/ifm-ai/search360/tree/4f9a64d47e50cdeec7ecf9c78d2fdfe4d4fcd594">search360</a>
is a retrieval service over <a
data-orig-href="https://github.com/ifm-ai/search360/blob/4f9a64d47e50cdeec7ecf9c78d2fdfe4d4fcd594/README.md#L209-L210" href="https://github.com/ifm-ai/search360/blob/4f9a64d47e50cdeec7ecf9c78d2fdfe4d4fcd594/README.md#L209-L210">1.6
billion passages</a> that <a
data-orig-href="https://github.com/ifm-ai/search360/blob/4f9a64d47e50cdeec7ecf9c78d2fdfe4d4fcd594/README.md#L16-L18" href="https://github.com/ifm-ai/search360/blob/4f9a64d47e50cdeec7ecf9c78d2fdfe4d4fcd594/README.md#L16-L18">"was
used to generate synthetic training data at billion-token scale"</a> for
K2 Horizon. <a
href="https://github.com/ifm-ai/LCQA/tree/1dee2447ef8c7ea5dddb7d0c2f4122f8bdcb5ede">LCQA</a>
and <a
href="https://github.com/ifm-ai/PRism-synthesis/tree/7fd0264ffe61cbeb309f580bb70fd55111ca996f">PRism-synthesis</a>
generate long-context questions and agentic code data. And the README of
<a
href="https://github.com/ifm-ai/process_entry/tree/a4745b10f5ce71e15d46021301e1ee717dc09bdc">process_entry</a>,
a validator for chat and agent trajectories, says <a
data-orig-href="https://github.com/ifm-ai/process_entry/blob/a4745b10f5ce71e15d46021301e1ee717dc09bdc/README.md#L3-L4" href="https://github.com/ifm-ai/process_entry/blob/a4745b10f5ce71e15d46021301e1ee717dc09bdc/README.md#L3-L4">"Every
conversation used in K2 Horizon mid- and post-training had to pass
it."</a> The generators behind most of the synthetic pretraining data
are not public, and neither is the mixture search.</p>
<p><strong>IFM's view of data.</strong> Mikhail Yurochkin, who leads the
LLM data team at IFM's Silicon Valley lab, set out its approach in a
talk, <a
href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf"><em>Diversity
First</em></a> (undated; the file name says April 9). Three of its
points bear on Marin:</p>
<ul>
<li><em>Mixture search paid little.</em> IFM fit a model that predicts
evaluation scores from mixture weights over sources defined by
WebOrganizer's labels, then sampled and refined candidate mixes. At 1.5B
parameters and 0.6T tokens, the searched mix scored about two points of
MMLU-CoT above a hand-tuned mix, reading the slide's chart. The slide's
title is <a
data-orig-href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=15" href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=15">"It
worked but improvement is Small"</a>, and it notes that the <a
data-orig-href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=15" href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=15">"data
mixing experiment required 375k GPU-hours"</a>. On BBH, the best of four
mixes at 1.5B was the worst at 7B (<a
data-orig-href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=11" href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=11">slide
11</a>).</li>
<li><em>Base-model benchmarks are the wrong target.</em> Reusing figures
from the <a href="https://arxiv.org/abs/2510.24397">APTBench</a> paper,
the talk shows base-model MMLU correlating weakly (r = 0.38) with
post-trained SWE-bench Verified scores. It proposes scoring mixes by
multiple-choice versions of post-training tasks, by pass@k, or by pass
rates after chat-formatted data is added to pretraining (<a
data-orig-href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=4" href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=4">slides
4–9</a>).</li>
<li><em>Diversity pays, if synthetic data is grounded in real data.</em>
The recap sets <a
data-orig-href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=16" href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=16">"Data
Mixing: many nuances; computationally expensive; small gains"</a>
against <a
data-orig-href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=16" href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=16">"Data
Diversity: steady and reliable gains"</a>. By IFM's compression measure,
reasoning text generated from seed queries grew nearly as repetitive as
web code; adding retrieval from an index of the pretraining corpus, and
randomness in the prompts, brought it close to high-quality web text (<a
data-orig-href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=24" href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=24">slide
24</a>).</li>
</ul>
<p>K2 Horizon applied this at scale: about <a
href="https://web.archive.org/web/20260913112353/https://ifm.ai/blog/k2/">"10
trillion synthetic tokens during pre-training"</a>, roughly half its
pretraining tokens, generated with <a
href="https://web.archive.org/web/20260913112353/https://ifm.ai/blog/k2/">"millions
of combinations of diversity knobs and context seeds, including
retrieval from an internally built search engine over the pre-training
web corpus"</a>. IFM has released five K2 Horizon datasets: <a
href="https://huggingface.co/datasets/IFM/TxT360-v2">TxT360-v2</a> (web
text, some with question-and-answer additions), <a
href="https://huggingface.co/datasets/IFM/Pretrain-Behaviors">Pretrain-Behaviors</a>,
<a
href="https://huggingface.co/datasets/IFM/Math-Reasoning">Math-Reasoning</a>,
<a
href="https://huggingface.co/datasets/IFM/Code-Reasoning">Code-Reasoning</a>
and <a
href="https://huggingface.co/datasets/IFM/SFT-Reasoning">SFT-Reasoning</a>.
By my rough estimate from sampled rows they hold about 13.5T tokens,
most of them synthetic. That is far more pretraining data than Marin has
released, but less than IFM's product page promises (<a
href="https://web.archive.org/web/20260906023511/https://ifm.ai/k2/">"We
provide the full pre-training corpus"</a>):</p>
<ul>
<li>The dataset cards name no generator models or prompts, so users
cannot check license terms inherited from the generators. They say only
that subsets <a
href="https://huggingface.co/datasets/IFM/TxT360-v2">"may have
undergone"</a> deduplication, and never mention decontamination.</li>
<li>No stage's mixture weights are published, although IFM published
exact weights for its previous model (<a
href="https://github.com/LLM360/k2v2_train">k2v2_train</a>).</li>
<li>The pretraining and mid-training datasets named in the model cards
are not public.</li>
<li>Plain web text covers only the Medium-High quality category, the one
subset that carries IFM's quality and topic labels. High-category pages
come with programmatic questions and answers appended, and the
<code>txt360-qa</code> subset, about half of TxT360-v2's tokens, appears
to be K2-V2-era data released again.</li>
<li>In a sampled slice of the <code>web-high-medium</code> subset, about
a tenth of the rows are labeled <code>clueweb</code>. Carnegie Mellon
distributes ClueWeb22 for research only, under signed license agreements
(<a
href="https://web.archive.org/web/20260925071610/https://lemurproject.org/clueweb22/obtain.php">terms</a>),
so users should confirm IFM's right to release those rows under
CC-BY-4.0.</li>
</ul>
<p><strong>Marin's Datakit curates open corpora.</strong> Will Held's <a
href="https://openathena.ai/blog/marin-data-pipeline-overview/">blog
post</a> describes the pipeline. It starts from about 25T tokens of
datasets that others have already collected and cleaned, plus one slice
of Common Crawl that Marin extracted itself. Normalization gives each
document a content hash. Exact and fuzzy deduplication then run across
sources, and a candidate duplicate is removed only if a longer kept
document contains at least 75% of its word 3-grams; the blog reports
12.5% of documents and 8.4% of tokens removed. Decontamination checked
18.7 billion documents and marked 259,960 (<a
href="https://github.com/marin-community/marin/pull/8327">#8327</a>).
The hero trains from a 23.1T-token store of the 200 topic-and-band
buckets.</p>
<p>Datakit has gaps. It has no language identification, rule-based
filters or corpus-wide PII removal. Its weekly Nemotron-CC canary has
failed every week since August 17; since September 7 the cause has been
a verification step that cannot get 1,000 workers ready in time (<a
href="https://github.com/marin-community/marin/issues/8948">#8948</a>,
<a
href="https://github.com/marin-community/marin/issues/9494">#9494</a>).
And the code behind the hero's distinctive choices is not on
<code>main</code>. The topic-clustering pull request was closed unmerged
(<a
href="https://github.com/marin-community/marin/pull/8208">#8208</a>),
the quality-scoring one is still open (<a
href="https://github.com/marin-community/marin/pull/8303">#8303</a>),
and both mixture-search ones were closed (<a
href="https://github.com/marin-community/marin/pull/7541">#7541</a>, <a
href="https://github.com/marin-community/marin/pull/8633">#8633</a>).
From <code>main</code>, Marin can rerun the same deduplication and
decontamination rules, but not the steps that assigned the hero's topics
and quality bands or chose its weights. Those survive only as published
artifacts: the <a
href="https://huggingface.co/marin-community/marin-data-mix-tools">topic
centroids and quality model</a> and the <a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/moe_hero_ep/harrier_mix_2026_08_18.py#L36-L49" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/moe_hero_ep/harrier_mix_2026_08_18.py#L36-L49">mixture
spec</a>. Marin does not rehost its corpus, because <a
href="https://discord.com/channels/1354881461060243556/1462895580064911522/1543322065484910652">"several
datasets we use forbid redistribution"</a>; its recipes download each
source from its creator. It does publish data it generates, such as its
<a
href="https://huggingface.co/datasets/marin-community/grug-moe-mix-swarm">proxy-run
results</a> and <a
href="https://huggingface.co/datasets/open-athena/Snowball-67B-A2B-RLVR1-Repro-Data">RL
reproduction data</a>.</p>
<p><strong>How Marin mixes.</strong> Marin searches with swarms of small
proxy MoE models. Each proxy trains on a scaled-down pool, repeating it
as often as the hero would repeat the full one; a regression fit to the
results proposes weights over the 200 buckets, and a ladder of larger
models checks them. The objective is BPB on held-out text and on the
answers of base-model benchmarks. Marin's results echo IFM's:</p>
<ul>
<li>In June, at 3e17–3e19 FLOPs, a curated mix tied proportional
sampling (1.14× compute-equivalent, 90% interval 0.93–1.41), while the
curated mix on Datakit's new data beat the old Nemotron-based mix 1.78×
(<a
data-orig-href="https://github.com/marin-community/marin/issues/6757#issuecomment-4847875950" href="https://github.com/marin-community/marin/issues/6757#issuecomment-4847875950">tie</a>,
<a
data-orig-href="https://github.com/marin-community/marin/issues/6757#issuecomment-4847173767" href="https://github.com/marin-community/marin/issues/6757#issuecomment-4847173767">new
data</a>). Will Held concluded that <a
href="https://discord.com/channels/1354881461060243556/1520123383138750595/1520140544926421136">"new
data is right now a far bigger accelerant"</a>.</li>
<li>The hero's August launch mix <a
href="https://discord.com/channels/1354881461060243556/1462895580064911522/1539387191216709673">"looked
to underperform our old mix"</a> on the pre-launch ladder.</li>
<li>A September re-mix beat the August mix 1.20× at the largest ladder
rung, in a single-seed comparison, and has fed the hero since step
108,000 on 2026-09-15 (<a
href="https://github.com/marin-community/marin/issues/9126">#9126</a>,
<a
data-orig-href="https://github.com/marin-community/marin/issues/8506#issuecomment-5684891157" href="https://github.com/marin-community/marin/issues/8506#issuecomment-5684891157">#8506</a>).
I found no run at scale that compares the final mix with proportional
sampling.</li>
</ul>
<p>The search is not cheap. By my estimate from Marin's records, the
June swarm of 840 proxy runs cost about 370,000 TPU-v4 chip-hours (<a
data-orig-href="https://github.com/marin-community/marin/issues/7067#issuecomment-4990793633" href="https://github.com/marin-community/marin/issues/7067#issuecomment-4990793633">#7067</a>),
comparable in scale to IFM's 375,000 GPU-hours, and the September swarm
27,000–30,000 H100-hours. The blog's own advice is modest: <a
href="https://openathena.ai/blog/marin-data-pipeline-overview/">"When in
doubt, using proportional sampling or UniMax will probably be better
than a poorly executed learned mixture."</a></p>
<p>The labs differ in their target and their synthetic share. Marin
still optimizes BPB, and its tests of post-training proxies were mixed:
base-model loss on math traces predicted post-RL accuracy across ten
models (r = −0.89) but not the gain from RL (R² = 0.33) (<a
href="https://github.com/marin-community/marin/issues/6096">#6096</a>).
About half of K2 Horizon's pretraining tokens were synthetic, while
Marin's synthetic data comes mostly from NVIDIA's Nemotron releases.
Marin's own research found that rephrasing web text gave 1.48× data
efficiency, and 1.80× when rephrasings were stitched into longer
documents, in a data-constrained setting (<a
href="https://github.com/marin-community/marin/issues/3905">#3905</a>).
Datakit has no stage that generates such data.</p>
<p><em>Verdict: Marin for code, IFM for released data. Datakit covers
deduplication, decontamination and tokenization with tests, and Marin
publishes its mixture weights, proxy-run data and ladder results, though
not the code that chose the hero's topics, bands and weights. IFM has
released trillions of tokens, most of them synthetic, but no mixture
weights and little of the code behind its newest data. Both labs found
that mixture search gains less than new or more diverse data.</em></p>
<h2 id="14-post-training">14. Post-training</h2>
<p>Both labs run RL in PyTorch with Megatron as the learner; all ten
open RL frameworks that Marin surveyed in June were PyTorch (<a
href="https://github.com/marin-community/marin/issues/6162">#6162</a>).
IFM runs RL with <a
href="https://github.com/ifm-ai/RL360/tree/5b548b6948e6e78af99307a38ee96057ea571faa">RL360</a>
and supervised fine-tuning (SFT) with xLLM, also PyTorch. Marin runs RL
with <a
href="https://github.com/marin-community/MarinSkyRL/tree/0cdccc9229f47c8eadff90aa2033a5b4604266aa">MarinSkyRL</a>
but fine-tunes in JAX.</p>
<table>
<thead>
<tr>
<th></th>
<th>IFM</th>
<th>Marin</th>
</tr>
</thead>
<tbody>
<tr>
<td>Recipe</td>
<td>Mid-training shifts toward reasoning and agentic data; then RL
experts, a weight merge, and 250–330B tokens of SFT at 512K (3.7B, 7B,
375B). The 0.9B, mid-trained only to 128K, has no SFT: on-policy
distillation from its experts follows the merge. The 32B and MoVA get
SFT only (269B tokens)</td>
<td>SFT (80% chat data, 20% pretraining replay, 262K context), then RL
with verifiable rewards across mixed domains; experts and distillation
planned</td>
</tr>
<tr>
<td>SFT code</td>
<td>xLLM (PyTorch): chat data in the pretraining loader, loss on
assistant tokens</td>
<td>JAX: a TPU-only copy of the Grug trainer on an unmerged branch, with
loss on all tokens, for the September Datakit cuts; Levanter with
assistant-only loss, from a launcher off <code>main</code>, for a later
cut</td>
</tr>
<tr>
<td>RL framework</td>
<td>Snapshot of IFM's public forks of Miles, Megatron-LM, SGLang, the
SMG router and Harbor</td>
<td>Hard fork of SkyRL; Megatron-Core only since 2026-09-25; Marin's
vLLM fork; a Harbor fork</td>
</tr>
<tr>
<td>Algorithms in the public recipe or configs</td>
<td>GRPO without standard-deviation scaling, DAPO-style clipping and
filtering, TIS, one-step-off asynchronous rollouts (<a
data-orig-href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/recipes/coding-overfit32.sbatch#L111-L124" href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/recipes/coding-overfit32.sbatch#L111-L124">recipe</a>)</td>
<td>GRPO and RLOO-N with PPO-style clipping, TIS, router replay,
bounded-staleness asynchronous training</td>
</tr>
<tr>
<td>Rewards</td>
<td>Harbor task tests; a 9,900-line verifier library, half of it bundled
benchmark code, that no launcher uses</td>
<td>Harbor tasks; skyrl-gym verifiers, including NVIDIA's Nemotron-Ultra
graders and generative reward model; LLM judges</td>
</tr>
<tr>
<td>Combining experts</td>
<td>Weight merges (ISO and RAM for the 3.7B and 7B, task arithmetic for
the 0.9B); code private, but experts, merged checkpoints and settings
released</td>
<td>No weight merging; multi-teacher on-policy distillation planned and
smoke-tested</td>
</tr>
<tr>
<td>Model support</td>
<td>K2 Horizon 7B recipe; 375B and MoVA code paths without recipes; no
export to Hugging Face format</td>
<td>Snowball 67B-A2B through a Megatron port of Grug; hero support in
open pull requests</td>
</tr>
<tr>
<td>Operations</td>
<td>Slurm and Docker; one recipe; no resume</td>
<td>Marin's Iris launcher; checkpoints streamed to S3; x86 hosts only on
<code>main</code>, so no GB200</td>
</tr>
<tr>
<td>Tests and CI</td>
<td>None for IFM's own code in the snapshot; the Miles fork keeps about
140 test files, but its test workflow is off</td>
<td>About 2,300 test functions; CPU tests on every pull request; three
nightly H100 lanes</td>
</tr>
<tr>
<td>Delivered</td>
<td>Six post-trained K2 Horizon models, 0.9B to 375B; RL experts for
four of them, including the 375B-A23B MoE</td>
<td>RL checkpoints of the 67B-A2B, released as research artifacts; first
official release pushed back a week from October 6</td>
</tr>
</tbody>
</table>
<p><strong>RL360 is a release snapshot.</strong> Its README sums up the
design: <a
data-orig-href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/README.md#L41" href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/README.md#L41">"Miles
manages rollouts and training, Megatron-LM updates the policy, SMG
routes requests to SGLang, and Harbor runs tasks and verifiers in
sandboxes"</a>. Outside its vendored components it holds 14,200 lines,
1.6% of the tree, and about 5,000 of those are bundled benchmark
verifiers such as IFEval and LiveBench. The rest is vendored from IFM's
public forks, without their Python tests. IFM's work inside those forks
is not in the 14,200: its Miles fork alone is 50 commits and about 9,600
added lines ahead of upstream (<a
href="https://github.com/LLM360/miles/compare/85fdb7e782388b9584150d72405d09cbd8ad15f4...de89a5ee12d026b52752268813b5c45ac3d9a132">compare</a>).
That work wires Harbor sandboxes into multi-turn rollouts, keeps sandbox
failures out of training, and adds K2 Horizon and MoVA support. The
Megatron-LM copy dates from February 2026 (megatron-core 0.16.0rc0).</p>
<p>The one recipe trains the 7B to overfit 32 coding tasks on 33
eight-GPU nodes: 64 GPUs train, 192 generate and one node runs services
(<a
data-orig-href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/README.md#L47" href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/README.md#L47">README</a>).
The README warns that <a
data-orig-href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/README.md#L165" href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/README.md#L165">"the
Docker build, checkpoint conversion, and full GPU run still need to be
tested together on the target cluster"</a>. Its <a
href="https://wandb.ai/mbzuai-llm/rl360-public/runs/wm7usgst">public
run</a>, on H200s, hit its 72-hour limit after 99 of 100 planned steps,
about 18,900 GPU-hours. Training reward rose from 0.37 to 0.90, while
the training GPUs sat idle 97% of each 40-minute step as agents worked
in sandboxes. The demo does not show IFM's production pace: the 7B's
shipped code experts logged median idle shares of 12–13% (<a
href="https://wandb.ai/llm360/K2-Horizon-7B/runs/m7tqeo5y">run</a>),
though its math and tool-use experts idled 63% and 88%.</p>
<p>RL360 is also not the exact code behind K2 Horizon's shipped experts.
The 3.7B's and 7B's math, code and STEM experts ran from an internal
RL360 checkout but logged metrics this snapshot cannot produce. Two of
the 7B's four experts <a
href="https://wandb.ai/llm360/K2-Horizon-7B/runs/4blppt0e">"were trained
by another team"</a>; the 0.9B's RL and distillation ran on slime (<a
href="https://wandb.ai/llm360/K2-Horizon-0.9B/runs/2y4evxdr">run</a>);
and the 375B's experts came from a harness named <code>fmp</code> (<a
href="https://wandb.ai/llm360/K2-Horizon-375B/runs/bot25bfe">run</a>).
The 7B's card describes an <a
href="https://huggingface.co/IFM/K2-Horizon-7B">"ISO merge on
self-attention and shared experts, RAM on the remaining weights"</a>,
but no merge code is public, and SFT ran on xLLM (<a
href="https://wandb.ai/llm360/K2-Horizon-7B/runs/wms33y11">run</a>). So
IFM's promise, <a
href="https://web.archive.org/web/20260906023511/https://ifm.ai/k2/">"We
make both our pre-training and post-training code available so you can
rerun or modify the training pipeline"</a>, holds only in part. Reading
the code also turns up three defects:</p>
<ul>
<li>The documented task file lacks a field that the agent code requires,
so rollouts set up as documented would abort (<a
data-orig-href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/src/agent360/harbor/miles/swe_agent_function.py#L110-L133" href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/src/agent360/harbor/miles/swe_agent_function.py#L110-L133">swe_agent_function.py</a>).</li>
<li>The snapshot cannot export a trained K2 Horizon policy to Hugging
Face format: the conversion tool was left out, and
<code>--save-hf</code> swallows its errors (<a
data-orig-href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/components/miles/miles/backends/megatron_utils/model.py#L751-L790" href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/components/miles/miles/backends/megatron_utils/model.py#L751-L790">model.py</a>).</li>
<li>The unused code verifier runs model-written code without a sandbox
by default (<a
data-orig-href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/src/miles360/reward/coder1/__init__.py#L15-L25" href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/src/miles360/reward/coder1/__init__.py#L15-L25">coder1</a>).</li>
</ul>
<p><strong>MarinSkyRL is a hard fork that Marin keeps
rewriting.</strong> Benjamin Feuer brought it from the OpenThoughts
project in June, and its policy reads: <a
data-orig-href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/AGENTS.md#L10-L11" href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/AGENTS.md#L10-L11">"This
is a <strong>hard snapshot</strong>: we own this tree and no upstream
sync or merge-back is planned"</a>. About 74% of its active Python lines
were last written in the Marin era, and 628 pull requests have merged
since mid-July, 81% of them Feuer's, so, like the hero's GB200 path, it
leans heavily on one engineer. It also changes fast. On 2026-09-25, <a
href="https://github.com/marin-community/MarinSkyRL/pull/776">#776</a>
deleted the FSDP, DeepSpeed and LoRA code, 36,000 lines, about half of
them tests. Megatron-Core is now the only trainer (<a
data-orig-href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/utils/utils.py#L482-L486" href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/utils/utils.py#L482-L486">utils.py</a>),
and its runtime installs only on x86 hosts, so GB200 support waits on an
open pull request (<a
href="https://github.com/marin-community/MarinSkyRL/pull/798">#798</a>).
On 2026-09-28, <a
href="https://github.com/marin-community/MarinSkyRL/pull/774">#774</a>
replaced the training loop. Three experiments on Marin's
<code>main</code> still request the deleted FSDP2 path (<a
href="https://github.com/marin-community/MarinSkyRL/issues/809">MarinSkyRL
#809</a>). RL trains with AdamW: the Megatron path is set up and tested
only for Adam and SGD (<a
data-orig-href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/distributed/megatron/optimizer.py#L26-L35" href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/distributed/megatron/optimizer.py#L26-L35">optimizer.py</a>),
and the Grug MuonH optimizer went with the FSDP code; MuonH ports are in
draft pull requests (<a
href="https://github.com/marin-community/MarinSkyRL/pull/820">#820</a>).
Both forks lag their upstreams. MarinSkyRL is 815 commits behind SkyRL,
and in August Marin judged a rebase intractable (<a
data-orig-href="https://github.com/marin-community/marin/issues/8048#issuecomment-5243952047" href="https://github.com/marin-community/marin/issues/8048#issuecomment-5243952047">"The
rebase is still intractable"</a>); IFM's Miles fork is 1,515 commits
behind Miles.</p>
<p>Rollouts run on Marin's vLLM fork. A weight sync that sends each
expert straight to the engine that serves it installs a Snowball update
in 0.51 s, against 19.9 s for a full broadcast (<a
data-orig-href="https://github.com/marin-community/marin/issues/8955#issuecomment-5634012950" href="https://github.com/marin-community/marin/issues/8955#issuecomment-5634012950">#8955</a>),
and EAGLE-3 draft models trained on Snowball's own rollouts made agentic
RL steps 1.42× faster (<a
href="https://github.com/marin-community/marin/issues/9114">#9114</a>).
The SGLang code is present but disabled. The registries hold the common
GRPO-family estimators and nine policy losses (<a
data-orig-href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/utils/algorithm_registry.py#L21-L27" href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/utils/algorithm_registry.py#L21-L27">estimators</a>,
<a
data-orig-href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/utils/algorithm_registry.py#L86-L96" href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/utils/algorithm_registry.py#L86-L96">losses</a>),
alongside TIS, router replay, bounded-staleness asynchronous training
(<a
data-orig-href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/config/ppo_base_config.yaml#L260-L275" href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/config/ppo_base_config.yaml#L260-L275">config</a>)
and on-policy distillation. There is no critic, and so no GAE-based PPO;
no reward-model or DPO training; and no standalone SFT trainer. About
2,250 CPU test functions run on every pull request, and the nightly H100
workflow, now three lanes, passed 62 of its 84 runs since July 15. But
no automated run covers asynchronous training, agentic RL on GPUs or
multi-node Megatron, and the nightly gate checks only that runs finish
with finite metrics.</p>
<p><strong>Crossing from JAX to PyTorch costs Marin.</strong> Snowball
was pretrained in JAX, so RL needs a PyTorch copy of Grug, checkpoint
conversion and parity tests. The first PyTorch port, written as a
correctness baseline, ran 27.3× slower than Levanter per update (<a
href="https://github.com/marin-community/marin/pull/7985">#7985</a>);
grouped expert kernels cut that to 1.38× on the same replay (<a
href="https://github.com/marin-community/MarinSkyRL/issues/259">MarinSkyRL
#259</a>). FSDP2 then failed a gate that required the trainer's token
probabilities to match vLLM's exactly, and Marin kept Megatron (<a
data-orig-href="https://github.com/marin-community/marin/issues/8955#issuecomment-5652989341" href="https://github.com/marin-community/marin/issues/8955#issuecomment-5652989341">#8955</a>).
Parity rests on a chain of tiny-model checks: a fixture generated by
Levanter, with hidden size 8 and four experts, checks the PyTorch model
at FP32 tolerances (<a
data-orig-href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/tests/grug_training_parity.py#L15-L19" href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/tests/grug_training_parity.py#L15-L19">test</a>),
and the PyTorch model then checks the Megatron port. No test compares
Levanter with Megatron at Snowball's full width. Marin chose SkyRL in
June partly because Miles and slime were <a
data-orig-href="https://github.com/marin-community/marin/issues/6162#issuecomment-4622595016" href="https://github.com/marin-community/marin/issues/6162#issuecomment-4622595016">"SGLang-only"</a>
and Marin wanted to keep vLLM. It deleted its in-process JAX RL engine
in August (<a
href="https://github.com/marin-community/marin/pull/8063">#8063</a>).
JAX learners now exist only as drafts (<a
href="https://github.com/marin-community/marin/pull/9050">#9050</a>, <a
href="https://github.com/marin-community/marin/pull/9144">#9144</a>),
though David Hall measured a JAX Grug learner at only about 4.6% slower
than Megatron's pipeline-parallel one on H100s (<a
href="https://github.com/marin-community/marin/issues/8970">#8970</a>).
IFM crosses a smaller gap: xLLM and Megatron-LM are both PyTorch, and
IFM's forks add grouped RMSNorm, xLLM's partial-RoPE layout and MoVA to
Megatron, and an xLLM checkpoint bridge to Miles.</p>
<p><strong>Results so far.</strong> Marin's release pick,
<code>Snowball-67B-A2B-10T-Mixed-RLVR-Sync-Step92</code>, came from
mixed-domain RL on 64 H100s per arm, after the September 11 SFT, and was
chosen on the same 26-benchmark panel that reports its scores (<a
href="https://github.com/marin-community/marin/issues/9359">#9359</a>,
<a
href="https://github.com/marin-community/marin/issues/9412">#9412</a>).
It scores 83.8% on MATH-500, 43.9% on AIME24, 43.1% on GPQA-Diamond and
43.7% on MMLU-Pro, but 3.0% on a 100-task SWE-bench Verified sample and
1.2% on Terminal-Bench 2. The agentic scores may understate the model:
the later September 21 SFT cut, in its default thinking mode, writes
tool calls that the Terminus-2 agent harness <a
data-orig-href="https://github.com/marin-community/marin/issues/9225#issuecomment-5786970615" href="https://github.com/marin-community/marin/issues/9225#issuecomment-5786970615">"cannot
parse"</a>, and with thinking off it solved 21 of the same 100 SWE-bench
tasks. The second RL stage was <a
href="https://github.com/marin-community/marin/issues/9359">"flat to
regressing"</a>, and pruning deleted the checkpoints at its pass@16
peaks. On an older lineage, SFT followed by 30 GRPO steps lifted the
SWE-bench sample from 14.5% to 30.7% (<a
data-orig-href="https://github.com/marin-community/marin/issues/9225#issuecomment-5777718831" href="https://github.com/marin-community/marin/issues/9225#issuecomment-5777718831">#9225</a>).
Like IFM's, Marin's results came from code that has since changed:
Step92 predates the new training loop, and the SWE-bench run used the
FSDP2 trainer that #776 deleted.</p>
<p><strong>Fine-tuning.</strong> Marin's SFT code is scattered. The
September 20 and 21 Datakit SFT cuts came from a copy of the Grug
trainer on an unmerged branch, which Will Held called <a
href="https://discord.com/channels/1354881461060243556/1551721992040882276/1552128612260388924">"TPU
specific still"</a> and which computes loss on every token, including
user and tool turns (<a
data-orig-href="https://github.com/marin-community/marin/blob/8a0ec6cf352792bf6da963c2a5977f312afdf571/experiments/grug_sft/special_token_lr.py#L94-L100" href="https://github.com/marin-community/marin/blob/8a0ec6cf352792bf6da963c2a5977f312afdf571/experiments/grug_sft/special_token_lr.py#L94-L100">special_token_lr.py</a>).
Asked why, Held answered <a
href="https://discord.com/channels/1354881461060243556/1552024155959066787/1552093562928103527">"No
well founded reason other than that masking them doesn't accelerate
training much"</a>, adding that past studies had found training on user
turns neutral to positive. A September 23 SFT used assistant-only loss
through Levanter, from a launcher not on <code>main</code> (<a
href="https://huggingface.co/open-athena/Grug-67B-A2B-GLM53-RLVR-SFT-2026.09.23">card</a>).
On <code>main</code>, a Grug SFT backend supports masked, packed
training (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/june_tpu_67b_a2b/moe/train.py#L443-L444" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/june_tpu_67b_a2b/moe/train.py#L443-L444">train.py</a>),
and a newer Levanter path for Snowball masks non-assistant tokens but
cannot pack sequences (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/models/snowball.py#L744-L745" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/models/snowball.py#L744-L745">snowball.py</a>);
its checked-in recipes run 4 and 10 steps. IFM's SFT ran on xLLM, whose
public code trains only on assistant tokens. No experiment on
<code>main</code> uses Levanter's DPO or LoRA code, and its LoRA wraps
only Haliax <code>Linear</code> layers, which Grug models lack (<a
data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/adaptor/lora.py#L158-L162" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/adaptor/lora.py#L158-L162">lora.py</a>).</p>
<p><strong>Merge or distill.</strong> No one has compared the two labs'
post-trained models under one protocol; Marin's panel excludes K2
Horizon because <a
href="https://github.com/marin-community/marin/issues/9412">"its suite
is incomplete"</a>. The recipes differ on a testable point. IFM trains
separate RL experts, merges their weights, and then repairs the merge
with long SFT or, at 0.9B, distillation. Benjamin Feuer, who leads
Marin's post-training, cautioned that <a
href="https://discord.com/channels/1354881461060243556/1374989195109466122/1551893133514645608">"the
evidence for most of these is pending"</a> before listing, among other
tips, <a
href="https://discord.com/channels/1354881461060243556/1374989195109466122/1551893133514645608">"Mixed-domain
RLVR &gt;&gt; hyperspecialized expert RL from a stability and learning
standpoint <em>but</em> not all regimes benefit equally ... agentic
benefits less"</a> and <a
href="https://discord.com/channels/1354881461060243556/1374989195109466122/1551893133514645608">"RL
after RL shows definite signs of catastrophic forgetting, even mixed
domain"</a>. Marin plans to fold experts into one model by multi-teacher
on-policy distillation (<a
href="https://github.com/marin-community/marin/issues/9250">#9250</a>);
a three-teacher smoke test on Snowball has merged (<a
href="https://github.com/marin-community/MarinSkyRL/pull/749">MarinSkyRL
#749</a>), but Marin has run no weight-merge experiment.</p>
<p><em>Verdict: Marin for RL code, IFM for results. MarinSkyRL is the
more complete and better-tested RL codebase, though it changes weekly
and leans on one engineer; RL360 shows IFM's agentic plumbing but not
the exact code that trained or merged the experts IFM shipped. For SFT
the order flips: IFM's ran on xLLM, whose public code masks
non-assistant tokens, while Marin's Datakit cuts came from an unmerged,
TPU-only branch that trains on every token. Both labs settled on
Megatron for the learner and Harbor for agentic tasks. A team starting
fresh should begin from upstream Miles, which serves rollouts only
through SGLang, or SkyRL, which also supports vLLM, and borrow from
these forks, which are tied to their labs' models and clusters.</em></p>
<h2 id="15-other-factors">15. Other factors</h2>
<ul>
<li><strong>Evaluation.</strong> xLLM ships downstream tasks, including
RULER and SCROLLS, but calls its evaluation <a
href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/eval/README.md">"preliminarily
designed for debugging"</a>. Levanter logs per-dataset losses and bits
per byte and integrates lm-eval-harness in process, but Marin scores
post-trained models by serving them on its vLLM fork and running
Evalchemy, lm-eval-harness and Harbor against the endpoint under a
written evaluation policy (<a
href="https://github.com/marin-community/marin/issues/9409">#9409</a>).</li>
<li><strong>Export and serving.</strong> xbridges converts only the
Transformer family, not Gekko; its own vLLM path pins vLLM 0.24.0 and
overwrites the model registry (<a
data-orig-href="https://github.com/ifm-ai/xbridges/blob/227f3fd26e30b024ce09bc7c556c64616c13276e/xbridges/vllm/add_xllm_to_vllm.sh#L4-L20" href="https://github.com/ifm-ai/xbridges/blob/227f3fd26e30b024ce09bc7c556c64616c13276e/xbridges/vllm/add_xllm_to_vllm.sh#L4-L20">script</a>),
though the released models also serve through stock vLLM and SGLang with
remote code. Classic Levanter round-trips about 15 Hugging Face
architectures; Grug exports a custom type served by Marin's vLLM
fork.</li>
<li><strong>Ecosystem and debugging.</strong> PyTorch offers eager
debugging, a larger hiring pool, and the Flash Linear Attention and
Transformer Engine libraries. JAX brings TPUs, but also rack-scale
compiles of about 20 minutes (<a
href="https://github.com/marin-community/marin/issues/8244">#8244</a>)
and compile-time deadlocks, such as ranks auto-tuning different kernel
sizes (<a
href="https://github.com/marin-community/marin/issues/9038">#9038</a>,
since fixed).</li>
<li><strong>People.</strong> xLLM's code has been public for one day,
with one listed contributor; other IFM authors appear in its test paths
and companion repositories. The data toolkit and RL360 went public the
same week, each essentially as one release commit, though RL360's
history survives in IFM's public forks. Marin merged 114 pull requests
in the last full week of September, from 13 people and 4 agent or bot
accounts, and MarinSkyRL merged another 49.</li>
<li><strong>License.</strong> xLLM, the data toolkit, RL360, Marin and
MarinSkyRL use Apache 2.0, though RL360's vendored Megatron-LM keeps
NVIDIA's license. IFM's K2 Horizon datasets use CC-BY-4.0 (TxT360-v2) or
Apache 2.0.</li>
</ul>
<h2 id="recommendation">Recommendation</h2>
<p><strong>For Marin: stay on Levanter.</strong> Switching would abandon
a JAX run 7.96T tokens into 18T, and the scheduler, evaluation and
serving built around it. xLLM lacks two things the hero depends on:
expert parallelism across a rack and Muon-family optimizers. For
frontier MoE, it would add only leading dense layers, dropless routing
and MoVA. Running both trainers would need checkpoint converters and
parity tests that neither side has; Marin's RL work already found the
JAX-to-PyTorch boundary <a
href="https://github.com/marin-community/marin/issues/7164">"orthogonal
to, and harder than, the TPU→GPU (hardware) boundary"</a>. Take four
things from xLLM:</p>
<ol type="1">
<li><strong>Run the dense benchmark.</strong> Train Llama3-8B at 8K on
the same H100s or H200s in both stacks, scored with one FLOP formula. If
xLLM holds above 43% and Grug trails, fix Grug's Hopper attention
forward before any Hopper-heavy run.</li>
<li><strong>Count causal attention in Levanter's MFU,</strong> or report
both conventions, so numbers compare across stacks and sequence
lengths.</li>
<li><strong>Study IFM's long-context schedule.</strong> It mid-trained
dense models through 32K, 128K and 512K with full attention, the
approach Marin plans for the hero's 65K and 262K phases. The data is
private; the schedule and throughput are public.</li>
<li><strong>Ask IFM about Gekko and MoVA.</strong> They are the two
ideas in xLLM that Marin has not tried, and no released model uses
Gekko. The groups already overlap: Marin serves K2 Horizon models (<a
href="https://github.com/marin-community/marin/pull/9314">#9314</a>),
and an IFM research scientist wrote on Marin's Discord that <a
href="https://discord.com/channels/1354881461060243556/1357057383830126652/1552365242665799831">"the
goals of the organizations seem to be quite aligned"</a>.</li>
</ol>
<p>Four more follow from the data and post-training comparison:</p>
<ol start="5" type="1">
<li><strong>Test IFM's released data.</strong> The five K2 Horizon
datasets hold about 13.5T tokens, mostly synthetic; TxT360-v2's
<code>web-high-medium</code> subset also carries IFM's quality and topic
labels. Check the ClueWeb licensing, pass the data through Datakit's
deduplication and decontamination, and let a swarm price it like any
other source. Opinion inside Marin is split: Mark Muchane planned to add
<a href="https://github.com/marin-community/marin/issues/8970">"whatever
from the MBZUAI datasets seems actually good"</a>, while David Hall,
after spot-checking one subset, was <a
href="https://discord.com/channels/1354881461060243556/1368297424086499359/1545190602994483270">"kinda
skeptical this will help a large model"</a>.</li>
<li><strong>Favor new data over further mixture search.</strong> New
data gave Marin 1.78× where a curated mix tied proportional sampling,
and IFM reports small gains from a 375,000-GPU-hour search. Grounded
synthetic data is the obvious next source: IFM trained on about 10T
tokens of it, and Marin's own rephrasing research found 1.5–1.8× data
efficiency, yet Datakit has no stage that generates data.</li>
<li><strong>Put the hero's data and SFT code on
<code>main</code>.</strong> Land the quality scorer (<a
href="https://github.com/marin-community/marin/pull/8303">#8303</a>),
restore the topic clustering and mixture search, and merge the release
SFT recipe with a deliberate choice of loss mask, so the next run can
regenerate its data and recipe rather than reload them.</li>
<li><strong>Compare weight merging with distillation.</strong> Merging
is cheap, though IFM follows each merge with long SFT or, at 0.9B,
distillation to repair it. IFM released its 7B experts, merged
checkpoint and merge settings (<a
href="https://wandb.ai/llm360/K2-Horizon-7B/runs/4blppt0e">run</a>), so
Marin can study the method on IFM's checkpoints before training experts
of its own.</li>
</ol>
<p><strong>For other teams,</strong> the hardware decides the
pretraining stack. On TPUs or GB200 racks, choose Levanter. On Hopper
clusters under Slurm, for dense models or long-context research, xLLM is
a credible start, proven to 7B. For frontier MoE on Hopper, xLLM has
probably trained a 375B-A23B model, but with expert parallelism confined
to one node and no published efficiency; compare it with TorchTitan and
Megatron-Core, which already have pipeline parallelism, FP8 and
cross-node expert parallelism. If you pretrain in JAX, budget for a
PyTorch port of each model for RL: Marin's first port ran 27× slower
until grouped kernels brought it to 1.4×, and its parity checks stop at
tiny models. For data, Datakit is the more complete and better-tested
pipeline, though beyond one machine it needs Marin's Iris scheduler;
from IFM, the released datasets are worth more than the toolkit. For RL,
start from upstream Miles or SkyRL and borrow from the forks: IFM's
Miles fork for Harbor-based agentic rollouts, and MarinSkyRL for router
replay, per-expert weight sync and its Megatron port of a JAX-trained
MoE.</p>
<h2 id="method-and-limits">Method and limits</h2>
<p>Thirty-one research agents read the seven main repositories and IFM's
related data repositories, searched Marin's records and gathered IFM's
public materials; six more reviewed the drafts adversarially, and I
checked the claims behind each verdict against code, logs and threads
myself. A follow-up pass examined the 375B's Hugging Face files and
W&amp;B project, and I read IFM's data talk slide by slide. Nothing ran
on accelerators; one agent ran the data toolkit's CPU stages on
synthetic data. IFM assembled its public W&amp;B projects after the fact
from private runs, so its GPU counts and hardware are inferred, and IFM
runs internal code beyond its public releases. My token counts for IFM's
datasets extrapolate from sampled rows and could be off by a fifth or
more. MarinSkyRL's own issues are outside Marin's mirrored records, so I
read them directly on GitHub.</p>
</div>
<footer class="provenance">
<p><em>Corpus: 2026-09-28 17:07 UTC &middot; Generated: 2026-09-29 08:01 UTC &middot; Viewing: <span id="viewing-time"></span></em></p>
<blockquote>
<p><em>Data: marinmirror — 225,653 chunks, built 2026-09-28 17:07 UTC ·
summaries through 2026-09-21_2026-09-27. Code: xLLM 889db38, xattn
0d6d73b, xbridges 227f3fd, pretraining-data-toolkit cd93e11, RL360
5b548b6, search360 4f9a64d, TxT360 07d98df, LCQA 1dee244,
PRism-synthesis 7fd0264, process_entry a4745b1, Marin main 4aa26ec
(release-SFT branch 8a0ec6c), MarinSkyRL 0cdccc9. IFM sources: launch
blog and product page (Wayback), Hugging Face model and dataset cards
and dataset statistics (375B at revision 82af3bc), public W&amp;B
projects (entities llm360 and mbzuai-llm), the Diversity First data
talk, arXiv 2601.06463, 2512.06201 and 2510.24397; TorchTitan benchmark
docs. Marin sources also include the Datakit blog post (openathena.ai)
and Marin's Hugging Face artifacts.</em></p>
<p><em>Query: "Do a detailed analysis of <a
href="https://github.com/ifm-ai/xllm">https://github.com/ifm-ai/xllm</a>,
a new library for specifying and training large LLMs, and compare it to
what we have in Levanter. Consider flexibility in specifying new
architectures, completeness of kinds of linear and other scalable
attention mechanisms implemented, FP8 and FP4 training, completeness of
kernels for different GPUs, ability to train on non-GPUs, dimensions of
parallelism implemented, ease of evolution by agents, code complexity,
test coverage, performance, and any other measure important for a team
building frontier LLMs who needs to choose between Levanter and xLLM.
Write a detailed report, then do an adversarial review, then publish a
gist." Follow-ups: "Does xLLM support sparse MoE architectures? Is there
evidence they used this codebase to train their largest model at <a
href="https://huggingface.co/IFM/K2-Horizon-375B-A23B">https://huggingface.co/IFM/K2-Horizon-375B-A23B</a>?"
and "Use the same protocol of detailed analysis and adversarial review
to update the report with data curation and mixing code from <a
href="https://github.com/ifm-ai/pretraining-data-toolkit">https://github.com/ifm-ai/pretraining-data-toolkit</a>
versus datakit from Marin as well as <a
href="https://github.com/ifm-ai/RL360">https://github.com/ifm-ai/RL360</a>
versus MarinSkyRL and other post-training code from Levanter and Marin.
Update the gist when you finish. Include information for IFM about data
curation and mixing from the presentation at <a
href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf">https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf</a>."</em></p>
<p><em>Sub-queries: xLLM architecture and config system · xLLM attention
and sequence mixers (xattn, FLA) · xLLM kernels, precision and hardware
· xLLM parallelism and checkpointing · xLLM tests, quality and benchmark
claims · xLLM external context (IFM, K2 Horizon, W&amp;B logs) ·
Levanter/Grug architecture flexibility · Levanter attention and linear
mixers · Levanter kernels, FP8 and accelerators · Levanter parallelism
and fault tolerance · Levanter tests, docs and agent affordances · Marin
measured MFU · Marin FP8/FP4 history · Marin attention and hybrid-mixer
experiments · Marin parallelism decisions · Marin kernel engineering and
maintenance · Marin TPU and non-NVIDIA use · Marin agent-driven
development · Marin JAX-vs-PyTorch design decisions · Marin long-context
work and mentions of xLLM/IFM · in-flight code branches · adversarial
reviews: xLLM fact-check, Levanter fact-check,
prose/completeness/fairness · follow-up: xLLM MoE support;
K2-Horizon-375B-A23B provenance (card, config, model code, W&amp;B runs,
run ID, Megatron option names) · round 2: IFM pretraining-data-toolkit
code · IFM data practice (talk, blog, released datasets, K2-V2 paper) ·
Marin Datakit code · Marin Datakit history and hero data · Marin
data-mixing methods and results · IFM RL360 code and W&amp;B runs ·
MarinSkyRL code · Marin RL infrastructure history · Marin post-training
recipes and results · Levanter post-training code · adversarial reviews
(round 2): IFM-side fact-check, Marin-side fact-check, prose,
completeness and fairness.</em></p>
</blockquote>
</footer>
</div>
</body>
</html>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment