Last active
September 29, 2026 08:01
-
-
Save hammer/a17688f9fcce5879ea7da426940ce2c4 to your computer and use it in GitHub Desktop.
IFM vs. Marin: pretraining, data and post-training stacks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| <!DOCTYPE html> | |
| <html lang="en"> | |
| <head> | |
| <meta charset="utf-8"> | |
| <meta name="viewport" content="width=device-width, initial-scale=1"> | |
| <title>IFM vs. Marin: pretraining, data and post-training stacks</title> | |
| <style> | |
| :root { | |
| --ink: #1a1a2e; --ink-secondary: #555770; --ink-faint: #8a8a9a; | |
| --accent: #8b2500; --accent-light: #c4530a; | |
| --surface: #faf9f6; --surface-raised: #f0eeea; --surface-code: #f5f3ef; | |
| --rule: #d4d0c8; --link: #8b2500; --link-hover: #c4530a; --ref-bg: #f7f5f1; | |
| --content-width: 650px; --sidenote-width: 230px; --sidenote-gap: 30px; | |
| } | |
| @media (prefers-color-scheme: dark) { | |
| :root:not([data-theme="light"]) { | |
| --ink: #d8d5cf; --ink-secondary: #9e9bab; --ink-faint: #6e6b7b; | |
| --accent: #d4764e; --accent-light: #e8956e; | |
| --surface: #1a1a24; --surface-raised: #242430; --surface-code: #20202c; | |
| --rule: #33333f; --link: #d4764e; --link-hover: #e8956e; --ref-bg: #1e1e2a; | |
| } | |
| } | |
| :root[data-theme="dark"] { | |
| --ink: #d8d5cf; --ink-secondary: #9e9bab; --ink-faint: #6e6b7b; | |
| --accent: #d4764e; --accent-light: #e8956e; | |
| --surface: #1a1a24; --surface-raised: #242430; --surface-code: #20202c; | |
| --rule: #33333f; --link: #d4764e; --link-hover: #e8956e; --ref-bg: #1e1e2a; | |
| } | |
| * { margin: 0; padding: 0; box-sizing: border-box; } | |
| body { | |
| background: var(--surface); color: var(--ink); | |
| font-family: 'Helvetica Neue', Helvetica, Arial, sans-serif; | |
| font-size: 16px; line-height: 1.7; | |
| -webkit-font-smoothing: antialiased; | |
| } | |
| .page { | |
| max-width: calc(var(--content-width) + var(--sidenote-width) + var(--sidenote-gap) + 80px); | |
| margin: 0 auto; padding: 3rem 40px 4rem; position: relative; | |
| } | |
| @media (max-width: 1060px) { | |
| .page { max-width: 100%; padding: 2rem 1.5rem 3rem; } | |
| } | |
| .content { max-width: var(--content-width); } | |
| .paper-header { max-width: var(--content-width); margin-bottom: 2.5rem; padding-bottom: 2rem; } | |
| .paper-header h1 { | |
| font-family: 'Palatino Linotype', Palatino, 'Book Antiqua', Georgia, serif; | |
| font-size: 2rem; font-weight: 600; line-height: 1.25; color: var(--ink); | |
| text-wrap: balance; margin-bottom: 0.6rem; letter-spacing: -0.01em; | |
| } | |
| .paper-meta { font-size: 0.875rem; color: var(--ink-secondary); line-height: 1.5; } | |
| .paper-meta .author { font-weight: 500; } | |
| .paper-meta .mumwelt-link { color: var(--ink-secondary); text-decoration: none; border-bottom: 1px dotted var(--ink-faint); } | |
| .paper-meta .mumwelt-link:hover { color: var(--accent); border-bottom-color: var(--accent); } | |
| .prompt-box { | |
| max-width: var(--content-width); margin-bottom: 2.5rem; position: relative; | |
| } | |
| .prompt-label { | |
| position: absolute; left: -0.6em; top: 50%; transform: translateY(-50%); | |
| font-family: 'Palatino Linotype', Palatino, 'Book Antiqua', Georgia, serif; | |
| font-size: 5rem; font-weight: 700; color: var(--ink); opacity: 0.08; | |
| line-height: 1; pointer-events: none; user-select: none; | |
| } | |
| .prompt-text { | |
| font-family: 'Palatino Linotype', Palatino, 'Book Antiqua', Georgia, serif; | |
| font-style: italic; font-size: 1.15rem; line-height: 1.55; color: var(--ink); | |
| position: relative; | |
| } | |
| .abstract { margin-bottom: 2.5rem; max-width: var(--content-width); } | |
| .abstract-label { | |
| font-size: 0.7rem; font-weight: 600; text-transform: uppercase; | |
| letter-spacing: 0.1em; color: var(--ink-secondary); margin-bottom: 0.5rem; | |
| } | |
| .abstract p { font-size: 0.92rem; line-height: 1.75; color: var(--ink); } | |
| h1, h2, h3 { | |
| font-family: 'Palatino Linotype', Palatino, 'Book Antiqua', Georgia, serif; | |
| font-weight: 600; color: var(--ink); text-wrap: balance; | |
| } | |
| h2 { font-size: 1.4rem; margin-top: 2.5rem; margin-bottom: 0.75rem; letter-spacing: -0.005em; } | |
| h3 { font-size: 1.1rem; margin-top: 1.75rem; margin-bottom: 0.5rem; } | |
| p { margin-bottom: 1rem; max-width: var(--content-width); } | |
| a { color: var(--link); text-decoration: none; border-bottom: 1px solid transparent; | |
| transition: border-color 0.15s, color 0.15s; } | |
| a:hover { color: var(--link-hover); border-bottom-color: var(--link-hover); } | |
| .date-label { cursor: default; border-bottom: 1px dotted var(--ink-faint); } | |
| strong { font-weight: 600; } | |
| .sidenote-checkbox { display: none; } | |
| .sidenote-toggle { display: none; } | |
| .sidenote { | |
| float: right; clear: right; width: var(--sidenote-width); | |
| margin-right: calc(-1 * (var(--sidenote-width) + var(--sidenote-gap))); | |
| margin-top: 0.2rem; margin-bottom: 1rem; | |
| font-size: 0.8rem; line-height: 1.5; color: var(--ink-secondary); | |
| } | |
| .sidenote-number { font-size: 0.7rem; font-weight: 600; color: var(--accent); margin-right: 0.3em; } | |
| @media (max-width: 1060px) { | |
| .sidenote-toggle { | |
| display: inline; cursor: pointer; color: var(--accent); | |
| font-size: 0.78rem; font-weight: 600; user-select: none; | |
| } | |
| .sidenote { | |
| float: none; display: none; width: 100%; margin: 0.4rem 0 0.75rem 0; | |
| font-size: 0.84rem; padding: 0.6rem 0.9rem; background: var(--surface-raised); | |
| border-radius: 4px; border-left: 2px solid var(--accent); | |
| } | |
| .sidenote-checkbox:checked + .sidenote { display: block; } | |
| .sidenote-number { display: none; } | |
| } | |
| blockquote { border-left: 2px solid var(--rule); padding-left: 1.25rem; | |
| margin: 1.25rem 0; color: var(--ink-secondary); font-style: italic; } | |
| code { font-family: 'SF Mono', Menlo, Consolas, monospace; font-size: 0.85em; | |
| background: var(--surface-code); padding: 0.15em 0.35em; border-radius: 3px; } | |
| pre { background: var(--surface-code); border: 1px solid var(--rule); border-radius: 4px; | |
| padding: 1rem 1.25rem; overflow-x: auto; margin: 1.25rem 0; max-width: var(--content-width); } | |
| pre code { background: none; padding: 0; font-size: 0.82rem; line-height: 1.6; } | |
| ul, ol { margin-bottom: 1rem; padding-left: 1.5rem; max-width: var(--content-width); } | |
| li { margin-bottom: 0.35rem; } | |
| li::marker { color: var(--ink-faint); } | |
| table { max-width: var(--content-width); border-collapse: collapse; width: 100%; | |
| margin: 1.25rem 0; font-variant-numeric: tabular-nums; font-size: 0.9rem; } | |
| thead { border-top: 2px solid var(--ink); border-bottom: 1px solid var(--ink); } | |
| th { font-weight: 600; padding: 0.3rem 0.75rem 0.35rem; text-align: left; | |
| line-height: 1.2; white-space: nowrap; font-size: 0.82rem; vertical-align: bottom; } | |
| td { padding: 0.35rem 0.75rem; border: none; vertical-align: top; } | |
| tbody { border-bottom: 1.5px solid var(--ink); } | |
| th:first-child, td:first-child { padding-left: 0; } | |
| th:last-child, td:last-child { padding-right: 0; } | |
| figure { margin: 2rem 0; max-width: var(--content-width); } | |
| figure img { width: 100%; border-radius: 3px; border: 1px solid var(--rule); } | |
| figcaption { font-size: 0.8rem; color: var(--ink-secondary); margin-top: 0.5rem; | |
| line-height: 1.5; font-style: italic; } | |
| footer.provenance { | |
| margin-top: 3rem; padding-top: 1rem; max-width: var(--content-width); | |
| font-size: 0.78rem; line-height: 1.6; color: var(--ink-faint); | |
| } | |
| footer.provenance p { margin-bottom: 0.3rem; } | |
| footer.provenance a { color: var(--ink-faint); } | |
| footer.provenance blockquote { margin: 0; padding: 0; border: none; color: inherit; } | |
| a[data-hover-title] { position: relative; } | |
| .hover-card { | |
| position: absolute; bottom: 100%; left: 50%; transform: translateX(-50%); | |
| width: 320px; max-width: 90vw; padding: 0.65rem 0.8rem; | |
| background: var(--surface-raised); border: 1px solid var(--rule); | |
| border-radius: 6px; box-shadow: 0 4px 12px rgba(0,0,0,0.1); | |
| font-size: 0.78rem; line-height: 1.45; color: var(--ink); | |
| pointer-events: none; z-index: 100; margin-bottom: 6px; | |
| opacity: 0; transition: opacity 0.12s; | |
| } | |
| a[data-hover-title]:hover .hover-card, | |
| a[data-hover-title]:focus .hover-card, | |
| a.cite:hover .hover-card { opacity: 1; } | |
| .hover-card .hc-title { font-weight: 600; margin-bottom: 0.2rem; } | |
| .hover-card .hc-meta { font-size: 0.72rem; color: var(--ink-faint); margin-bottom: 0.25rem; } | |
| .hover-card .hc-status { | |
| display: inline-block; font-size: 0.65rem; font-weight: 600; | |
| text-transform: uppercase; letter-spacing: 0.04em; | |
| padding: 0.1em 0.4em; border-radius: 3px; margin-right: 0.4em; | |
| } | |
| .hc-status-open { background: #fff3e0; color: #e65100; } | |
| .hc-status-closed, .hc-status-merged { background: #e8f5e9; color: #2e7d32; } | |
| @media (prefers-color-scheme: dark) { | |
| :root:not([data-theme="light"]) .hover-card { box-shadow: 0 4px 12px rgba(0,0,0,0.3); } | |
| :root:not([data-theme="light"]) .hc-status-open { background: #3a2a10; color: #ffb74d; } | |
| :root:not([data-theme="light"]) .hc-status-closed, | |
| :root:not([data-theme="light"]) .hc-status-merged { background: #1b3a1e; color: #66bb6a; } | |
| } | |
| :root[data-theme="dark"] .hover-card { box-shadow: 0 4px 12px rgba(0,0,0,0.3); } | |
| :root[data-theme="dark"] .hc-status-open { background: #3a2a10; color: #ffb74d; } | |
| :root[data-theme="dark"] .hc-status-closed, | |
| :root[data-theme="dark"] .hc-status-merged { background: #1b3a1e; color: #66bb6a; } | |
| .hover-card .hc-desc { color: var(--ink-secondary); } | |
| .cite { | |
| font-size: 0.72rem; vertical-align: super; line-height: 0; | |
| color: var(--accent); font-weight: 600; text-decoration: none; | |
| border-bottom: none !important; position: relative; | |
| } | |
| .cite:hover { color: var(--link-hover); } | |
| .references { margin-top: 3rem; max-width: var(--content-width); } | |
| .references h2 { font-size: 1.15rem; margin-bottom: 1rem; } | |
| .ref-list { list-style: none; padding: 0; counter-reset: ref; } | |
| .ref-list li { | |
| counter-increment: ref; display: flex; align-items: baseline; | |
| gap: 0.5em; font-size: 0.82rem; line-height: 1.55; | |
| margin-bottom: 0.4rem; color: var(--ink-secondary); | |
| } | |
| .ref-list li::before { | |
| content: "[" counter(ref) "]"; flex-shrink: 0; | |
| font-variant-numeric: tabular-nums; color: var(--ink-faint); | |
| font-size: 0.78rem; min-width: 2.2em; | |
| } | |
| .ref-list .ref-body { flex: 1; min-width: 0; } | |
| .ref-list .ref-title { font-weight: 500; color: var(--ink); } | |
| .ref-list .ref-url { | |
| font-family: 'SF Mono', Menlo, Consolas, monospace; font-size: 0.75rem; | |
| color: var(--ink-faint); word-break: break-all; margin-left: 0.4em; | |
| } | |
| .ref-list .ref-url a { color: var(--ink-faint); border-bottom: none; } | |
| .ref-list .ref-url a:hover { color: var(--link-hover); } | |
| .ref-back { | |
| color: var(--accent); text-decoration: none; border-bottom: none !important; | |
| margin-left: 0.3em; font-size: 0.78rem; | |
| } | |
| .ref-back:hover { color: var(--link-hover); } | |
| .status { | |
| display: inline-block; font-size: 0.65rem; font-weight: 600; | |
| text-transform: uppercase; letter-spacing: 0.04em; | |
| padding: 0.15em 0.5em; border-radius: 3px; vertical-align: middle; | |
| } | |
| .status-done { background: #e8f5e9; color: #2e7d32; } | |
| .status-open { background: #fff3e0; color: #e65100; } | |
| .status-blocked { background: #fce4ec; color: #c62828; } | |
| @media (prefers-color-scheme: dark) { | |
| :root:not([data-theme="light"]) .status-done { background: #1b3a1e; color: #66bb6a; } | |
| :root:not([data-theme="light"]) .status-open { background: #3a2a10; color: #ffb74d; } | |
| :root:not([data-theme="light"]) .status-blocked { background: #3a1520; color: #ef9a9a; } | |
| } | |
| :root[data-theme="dark"] .status-done { background: #1b3a1e; color: #66bb6a; } | |
| :root[data-theme="dark"] .status-open { background: #3a2a10; color: #ffb74d; } | |
| :root[data-theme="dark"] .status-blocked { background: #3a1520; color: #ef9a9a; } | |
| .cite-card { | |
| position: absolute; bottom: 100%; left: 50%; transform: translateX(-50%); | |
| width: 360px; max-width: 90vw; padding: 0.65rem 0.8rem; | |
| background: var(--surface-raised); border: 1px solid var(--rule); | |
| border-radius: 6px; box-shadow: 0 4px 12px rgba(0,0,0,0.1); | |
| font-size: 0.78rem; line-height: 1.45; color: var(--ink); | |
| pointer-events: none; z-index: 100; margin-bottom: 6px; | |
| opacity: 0; transition: opacity 0.12s; | |
| font-weight: 400; vertical-align: baseline; text-align: left; | |
| } | |
| a.cite:hover .cite-card { opacity: 1; } | |
| @media (prefers-color-scheme: dark) { | |
| :root:not([data-theme="light"]) .cite-card { box-shadow: 0 4px 12px rgba(0,0,0,0.3); } | |
| } | |
| :root[data-theme="dark"] .cite-card { box-shadow: 0 4px 12px rgba(0,0,0,0.3); } | |
| @media print { | |
| body { font-size: 11pt; } | |
| .page { max-width: 100%; padding: 0; } | |
| .sidenote { float: right; width: 180px; margin-right: -210px; } | |
| .sidenote-toggle { display: none; } | |
| a { color: inherit; border-bottom: none; } | |
| .hover-card { display: none; } | |
| .cite-card { display: none; } | |
| } | |
| .katex{font:normal 1.21em KaTeX_Main,Times New Roman,serif;line-height:1.2;position:relative;text-indent:0;text-rendering:auto}.katex *{-ms-high-contrast-adjust:none!important;border-color:currentColor}.katex .katex-version:after{content:"0.18.9"}.katex .katex-mathml{border:0;-webkit-clip-path:inset(50%);clip-path:inset(50%);height:1px;overflow:hidden;padding:0;position:absolute;width:1px}.katex .katex-html>.katex-newline{display:block}.katex .katex-base{position:relative;white-space:nowrap;width:-webkit-min-content;width:-moz-min-content;width:min-content}.katex .katex-base,.katex .katex-strut{display:inline-block}.katex .textbf{font-weight:700}.katex .textit{font-style:italic}.katex .textrm{font-family:KaTeX_Main}.katex .textsf{font-family:KaTeX_SansSerif}.katex .texttt{font-family:KaTeX_Typewriter}.katex .mathnormal{font-family:KaTeX_Math;font-style:italic}.katex .mathit{font-family:KaTeX_Main;font-style:italic}.katex .mathrm{font-style:normal}.katex .mathbf{font-family:KaTeX_Main;font-weight:700}.katex .boldsymbol{font-family:KaTeX_Math;font-style:italic;font-weight:700}.katex .amsrm,.katex .mathbb,.katex .textbb{font-family:KaTeX_AMS}.katex .mathcal{font-family:KaTeX_Caligraphic}.katex .mathfrak,.katex .textfrak{font-family:KaTeX_Fraktur}.katex .mathboldfrak,.katex .textboldfrak{font-family:KaTeX_Fraktur;font-weight:700}.katex .mathtt{font-family:KaTeX_Typewriter}.katex .mathscr,.katex .textscr{font-family:KaTeX_Script}.katex .mathsf,.katex .textsf{font-family:KaTeX_SansSerif}.katex .mathboldsf,.katex .textboldsf{font-family:KaTeX_SansSerif;font-weight:700}.katex .mathitsf,.katex .mathsfit,.katex .textitsf{font-family:KaTeX_SansSerif;font-style:italic}.katex .mainrm{font-family:KaTeX_Main;font-style:normal}.katex .vlist-t{border-collapse:collapse;display:inline-table;table-layout:fixed}.katex .vlist-r{display:table-row}.katex .vlist{display:table-cell;position:relative;vertical-align:bottom}.katex .vlist>span{display:block;height:0;position:relative}.katex .vlist>span>span{display:inline-block}.katex .vlist>span>.pstrut{overflow:hidden;width:0}.katex .vlist-t2{margin-right:-2px}.katex .vlist-s{display:table-cell;font-size:1px;min-width:2px;vertical-align:bottom;width:2px}.katex .katex-vbox{align-items:baseline;display:inline-flex;flex-direction:column}.katex .katex-thinbox{display:inline-flex;flex-direction:row;max-width:0;width:0}.katex .msupsub{text-align:left}.katex .mfrac>span>span{text-align:center}.katex .mfrac .frac-line{border-bottom-style:solid;display:inline-block;width:100%}.katex .katex-hdashline,.katex .katex-hline,.katex .katex-overline .overline-line,.katex .katex-rule,.katex .katex-underline .underline-line,.katex .mfrac .frac-line{min-height:1px}.katex .mspace{display:inline-block}.katex .katex-smash{display:inline;line-height:0}.katex .clap,.katex .llap,.katex .rlap{position:relative;width:0}.katex .clap>.katex-inner,.katex .llap>.katex-inner,.katex .rlap>.katex-inner{position:absolute}.katex .clap>.katex-fix,.katex .llap>.katex-fix,.katex .rlap>.katex-fix{display:inline-block}.katex .llap>.katex-inner{right:0}.katex .clap>.katex-inner,.katex .rlap>.katex-inner{left:0}.katex .clap>.katex-inner>span{margin-left:-50%;margin-right:50%}.katex .katex-rule{border:0 solid;display:inline-block;position:relative}.katex .katex-hline,.katex .katex-overline .overline-line,.katex .katex-underline .underline-line{border-bottom-style:solid;display:inline-block;width:100%}.katex .katex-hdashline{border-bottom-style:dashed;display:inline-block;width:100%}.katex .sqrt>.katex-root{margin-left:.2777777778em;margin-right:-.5555555556em}.katex .fontsize-ensurer.reset-size1.size1,.katex .katex-sizing.reset-size1.size1{font-size:1em}.katex .fontsize-ensurer.reset-size1.size2,.katex .katex-sizing.reset-size1.size2{font-size:1.2em}.katex .fontsize-ensurer.reset-size1.size3,.katex .katex-sizing.reset-size1.size3{font-size:1.4em}.katex .fontsize-ensurer.reset-size1.size4,.katex .katex-sizing.reset-size1.size4{font-size:1.6em}.katex .fontsize-ensurer.reset-size1.size5,.katex .katex-sizing.reset-size1.size5{font-size:1.8em}.katex .fontsize-ensurer.reset-size1.size6,.katex .katex-sizing.reset-size1.size6{font-size:2em}.katex .fontsize-ensurer.reset-size1.size7,.katex .katex-sizing.reset-size1.size7{font-size:2.4em}.katex .fontsize-ensurer.reset-size1.size8,.katex .katex-sizing.reset-size1.size8{font-size:2.88em}.katex .fontsize-ensurer.reset-size1.size9,.katex .katex-sizing.reset-size1.size9{font-size:3.456em}.katex .fontsize-ensurer.reset-size1.size10,.katex .katex-sizing.reset-size1.size10{font-size:4.148em}.katex .fontsize-ensurer.reset-size1.size11,.katex .katex-sizing.reset-size1.size11{font-size:4.976em}.katex .fontsize-ensurer.reset-size2.size1,.katex .katex-sizing.reset-size2.size1{font-size:.8333333333em}.katex .fontsize-ensurer.reset-size2.size2,.katex .katex-sizing.reset-size2.size2{font-size:1em}.katex .fontsize-ensurer.reset-size2.size3,.katex .katex-sizing.reset-size2.size3{font-size:1.1666666667em}.katex .fontsize-ensurer.reset-size2.size4,.katex .katex-sizing.reset-size2.size4{font-size:1.3333333333em}.katex .fontsize-ensurer.reset-size2.size5,.katex .katex-sizing.reset-size2.size5{font-size:1.5em}.katex .fontsize-ensurer.reset-size2.size6,.katex .katex-sizing.reset-size2.size6{font-size:1.6666666667em}.katex .fontsize-ensurer.reset-size2.size7,.katex .katex-sizing.reset-size2.size7{font-size:2em}.katex .fontsize-ensurer.reset-size2.size8,.katex .katex-sizing.reset-size2.size8{font-size:2.4em}.katex .fontsize-ensurer.reset-size2.size9,.katex .katex-sizing.reset-size2.size9{font-size:2.88em}.katex .fontsize-ensurer.reset-size2.size10,.katex .katex-sizing.reset-size2.size10{font-size:3.4566666667em}.katex .fontsize-ensurer.reset-size2.size11,.katex .katex-sizing.reset-size2.size11{font-size:4.1466666667em}.katex .fontsize-ensurer.reset-size3.size1,.katex .katex-sizing.reset-size3.size1{font-size:.7142857143em}.katex .fontsize-ensurer.reset-size3.size2,.katex .katex-sizing.reset-size3.size2{font-size:.8571428571em}.katex .fontsize-ensurer.reset-size3.size3,.katex .katex-sizing.reset-size3.size3{font-size:1em}.katex .fontsize-ensurer.reset-size3.size4,.katex .katex-sizing.reset-size3.size4{font-size:1.1428571429em}.katex .fontsize-ensurer.reset-size3.size5,.katex .katex-sizing.reset-size3.size5{font-size:1.2857142857em}.katex .fontsize-ensurer.reset-size3.size6,.katex .katex-sizing.reset-size3.size6{font-size:1.4285714286em}.katex .fontsize-ensurer.reset-size3.size7,.katex .katex-sizing.reset-size3.size7{font-size:1.7142857143em}.katex .fontsize-ensurer.reset-size3.size8,.katex .katex-sizing.reset-size3.size8{font-size:2.0571428571em}.katex .fontsize-ensurer.reset-size3.size9,.katex .katex-sizing.reset-size3.size9{font-size:2.4685714286em}.katex .fontsize-ensurer.reset-size3.size10,.katex .katex-sizing.reset-size3.size10{font-size:2.9628571429em}.katex .fontsize-ensurer.reset-size3.size11,.katex .katex-sizing.reset-size3.size11{font-size:3.5542857143em}.katex .fontsize-ensurer.reset-size4.size1,.katex .katex-sizing.reset-size4.size1{font-size:.625em}.katex .fontsize-ensurer.reset-size4.size2,.katex .katex-sizing.reset-size4.size2{font-size:.75em}.katex .fontsize-ensurer.reset-size4.size3,.katex .katex-sizing.reset-size4.size3{font-size:.875em}.katex .fontsize-ensurer.reset-size4.size4,.katex .katex-sizing.reset-size4.size4{font-size:1em}.katex .fontsize-ensurer.reset-size4.size5,.katex .katex-sizing.reset-size4.size5{font-size:1.125em}.katex .fontsize-ensurer.reset-size4.size6,.katex .katex-sizing.reset-size4.size6{font-size:1.25em}.katex .fontsize-ensurer.reset-size4.size7,.katex .katex-sizing.reset-size4.size7{font-size:1.5em}.katex .fontsize-ensurer.reset-size4.size8,.katex .katex-sizing.reset-size4.size8{font-size:1.8em}.katex .fontsize-ensurer.reset-size4.size9,.katex .katex-sizing.reset-size4.size9{font-size:2.16em}.katex .fontsize-ensurer.reset-size4.size10,.katex .katex-sizing.reset-size4.size10{font-size:2.5925em}.katex .fontsize-ensurer.reset-size4.size11,.katex .katex-sizing.reset-size4.size11{font-size:3.11em}.katex .fontsize-ensurer.reset-size5.size1,.katex .katex-sizing.reset-size5.size1{font-size:.5555555556em}.katex .fontsize-ensurer.reset-size5.size2,.katex .katex-sizing.reset-size5.size2{font-size:.6666666667em}.katex .fontsize-ensurer.reset-size5.size3,.katex .katex-sizing.reset-size5.size3{font-size:.7777777778em}.katex .fontsize-ensurer.reset-size5.size4,.katex .katex-sizing.reset-size5.size4{font-size:.8888888889em}.katex .fontsize-ensurer.reset-size5.size5,.katex .katex-sizing.reset-size5.size5{font-size:1em}.katex .fontsize-ensurer.reset-size5.size6,.katex .katex-sizing.reset-size5.size6{font-size:1.1111111111em}.katex .fontsize-ensurer.reset-size5.size7,.katex .katex-sizing.reset-size5.size7{font-size:1.3333333333em}.katex .fontsize-ensurer.reset-size5.size8,.katex .katex-sizing.reset-size5.size8{font-size:1.6em}.katex .fontsize-ensurer.reset-size5.size9,.katex .katex-sizing.reset-size5.size9{font-size:1.92em}.katex .fontsize-ensurer.reset-size5.size10,.katex .katex-sizing.reset-size5.size10{font-size:2.3044444444em}.katex .fontsize-ensurer.reset-size5.size11,.katex .katex-sizing.reset-size5.size11{font-size:2.7644444444em}.katex .fontsize-ensurer.reset-size6.size1,.katex .katex-sizing.reset-size6.size1{font-size:.5em}.katex .fontsize-ensurer.reset-size6.size2,.katex .katex-sizing.reset-size6.size2{font-size:.6em}.katex .fontsize-ensurer.reset-size6.size3,.katex .katex-sizing.reset-size6.size3{font-size:.7em}.katex .fontsize-ensurer.reset-size6.size4,.katex .katex-sizing.reset-size6.size4{font-size:.8em}.katex .fontsize-ensurer.reset-size6.size5,.katex .katex-sizing.reset-size6.size5{font-size:.9em}.katex .fontsize-ensurer.reset-size6.size6,.katex .katex-sizing.reset-size6.size6{font-size:1em}.katex .fontsize-ensurer.reset-size6.size7,.katex .katex-sizing.reset-size6.size7{font-size:1.2em}.katex .fontsize-ensurer.reset-size6.size8,.katex .katex-sizing.reset-size6.size8{font-size:1.44em}.katex .fontsize-ensurer.reset-size6.size9,.katex .katex-sizing.reset-size6.size9{font-size:1.728em}.katex .fontsize-ensurer.reset-size6.size10,.katex .katex-sizing.reset-size6.size10{font-size:2.074em}.katex .fontsize-ensurer.reset-size6.size11,.katex .katex-sizing.reset-size6.size11{font-size:2.488em}.katex .fontsize-ensurer.reset-size7.size1,.katex .katex-sizing.reset-size7.size1{font-size:.4166666667em}.katex .fontsize-ensurer.reset-size7.size2,.katex .katex-sizing.reset-size7.size2{font-size:.5em}.katex .fontsize-ensurer.reset-size7.size3,.katex .katex-sizing.reset-size7.size3{font-size:.5833333333em}.katex .fontsize-ensurer.reset-size7.size4,.katex .katex-sizing.reset-size7.size4{font-size:.6666666667em}.katex .fontsize-ensurer.reset-size7.size5,.katex .katex-sizing.reset-size7.size5{font-size:.75em}.katex .fontsize-ensurer.reset-size7.size6,.katex .katex-sizing.reset-size7.size6{font-size:.8333333333em}.katex .fontsize-ensurer.reset-size7.size7,.katex .katex-sizing.reset-size7.size7{font-size:1em}.katex .fontsize-ensurer.reset-size7.size8,.katex .katex-sizing.reset-size7.size8{font-size:1.2em}.katex .fontsize-ensurer.reset-size7.size9,.katex .katex-sizing.reset-size7.size9{font-size:1.44em}.katex .fontsize-ensurer.reset-size7.size10,.katex .katex-sizing.reset-size7.size10{font-size:1.7283333333em}.katex .fontsize-ensurer.reset-size7.size11,.katex .katex-sizing.reset-size7.size11{font-size:2.0733333333em}.katex .fontsize-ensurer.reset-size8.size1,.katex .katex-sizing.reset-size8.size1{font-size:.3472222222em}.katex .fontsize-ensurer.reset-size8.size2,.katex .katex-sizing.reset-size8.size2{font-size:.4166666667em}.katex .fontsize-ensurer.reset-size8.size3,.katex .katex-sizing.reset-size8.size3{font-size:.4861111111em}.katex .fontsize-ensurer.reset-size8.size4,.katex .katex-sizing.reset-size8.size4{font-size:.5555555556em}.katex .fontsize-ensurer.reset-size8.size5,.katex .katex-sizing.reset-size8.size5{font-size:.625em}.katex .fontsize-ensurer.reset-size8.size6,.katex .katex-sizing.reset-size8.size6{font-size:.6944444444em}.katex .fontsize-ensurer.reset-size8.size7,.katex .katex-sizing.reset-size8.size7{font-size:.8333333333em}.katex .fontsize-ensurer.reset-size8.size8,.katex .katex-sizing.reset-size8.size8{font-size:1em}.katex .fontsize-ensurer.reset-size8.size9,.katex .katex-sizing.reset-size8.size9{font-size:1.2em}.katex .fontsize-ensurer.reset-size8.size10,.katex .katex-sizing.reset-size8.size10{font-size:1.4402777778em}.katex .fontsize-ensurer.reset-size8.size11,.katex .katex-sizing.reset-size8.size11{font-size:1.7277777778em}.katex .fontsize-ensurer.reset-size9.size1,.katex .katex-sizing.reset-size9.size1{font-size:.2893518519em}.katex .fontsize-ensurer.reset-size9.size2,.katex .katex-sizing.reset-size9.size2{font-size:.3472222222em}.katex .fontsize-ensurer.reset-size9.size3,.katex .katex-sizing.reset-size9.size3{font-size:.4050925926em}.katex .fontsize-ensurer.reset-size9.size4,.katex .katex-sizing.reset-size9.size4{font-size:.462962963em}.katex .fontsize-ensurer.reset-size9.size5,.katex .katex-sizing.reset-size9.size5{font-size:.5208333333em}.katex .fontsize-ensurer.reset-size9.size6,.katex .katex-sizing.reset-size9.size6{font-size:.5787037037em}.katex .fontsize-ensurer.reset-size9.size7,.katex .katex-sizing.reset-size9.size7{font-size:.6944444444em}.katex .fontsize-ensurer.reset-size9.size8,.katex .katex-sizing.reset-size9.size8{font-size:.8333333333em}.katex .fontsize-ensurer.reset-size9.size9,.katex .katex-sizing.reset-size9.size9{font-size:1em}.katex .fontsize-ensurer.reset-size9.size10,.katex .katex-sizing.reset-size9.size10{font-size:1.2002314815em}.katex .fontsize-ensurer.reset-size9.size11,.katex .katex-sizing.reset-size9.size11{font-size:1.4398148148em}.katex .fontsize-ensurer.reset-size10.size1,.katex .katex-sizing.reset-size10.size1{font-size:.2410800386em}.katex .fontsize-ensurer.reset-size10.size2,.katex .katex-sizing.reset-size10.size2{font-size:.2892960463em}.katex .fontsize-ensurer.reset-size10.size3,.katex .katex-sizing.reset-size10.size3{font-size:.337512054em}.katex .fontsize-ensurer.reset-size10.size4,.katex .katex-sizing.reset-size10.size4{font-size:.3857280617em}.katex .fontsize-ensurer.reset-size10.size5,.katex .katex-sizing.reset-size10.size5{font-size:.4339440694em}.katex .fontsize-ensurer.reset-size10.size6,.katex .katex-sizing.reset-size10.size6{font-size:.4821600771em}.katex .fontsize-ensurer.reset-size10.size7,.katex .katex-sizing.reset-size10.size7{font-size:.5785920926em}.katex .fontsize-ensurer.reset-size10.size8,.katex .katex-sizing.reset-size10.size8{font-size:.6943105111em}.katex .fontsize-ensurer.reset-size10.size9,.katex .katex-sizing.reset-size10.size9{font-size:.8331726133em}.katex .fontsize-ensurer.reset-size10.size10,.katex .katex-sizing.reset-size10.size10{font-size:1em}.katex .fontsize-ensurer.reset-size10.size11,.katex .katex-sizing.reset-size10.size11{font-size:1.1996142719em}.katex .fontsize-ensurer.reset-size11.size1,.katex .katex-sizing.reset-size11.size1{font-size:.2009646302em}.katex .fontsize-ensurer.reset-size11.size2,.katex .katex-sizing.reset-size11.size2{font-size:.2411575563em}.katex .fontsize-ensurer.reset-size11.size3,.katex .katex-sizing.reset-size11.size3{font-size:.2813504823em}.katex .fontsize-ensurer.reset-size11.size4,.katex .katex-sizing.reset-size11.size4{font-size:.3215434084em}.katex .fontsize-ensurer.reset-size11.size5,.katex .katex-sizing.reset-size11.size5{font-size:.3617363344em}.katex .fontsize-ensurer.reset-size11.size6,.katex .katex-sizing.reset-size11.size6{font-size:.4019292605em}.katex .fontsize-ensurer.reset-size11.size7,.katex .katex-sizing.reset-size11.size7{font-size:.4823151125em}.katex .fontsize-ensurer.reset-size11.size8,.katex .katex-sizing.reset-size11.size8{font-size:.578778135em}.katex .fontsize-ensurer.reset-size11.size9,.katex .katex-sizing.reset-size11.size9{font-size:.6945337621em}.katex .fontsize-ensurer.reset-size11.size10,.katex .katex-sizing.reset-size11.size10{font-size:.8336012862em}.katex .fontsize-ensurer.reset-size11.size11,.katex .katex-sizing.reset-size11.size11{font-size:1em}.katex .delimsizing.size1{font-family:KaTeX_Size1}.katex .delimsizing.size2{font-family:KaTeX_Size2}.katex .delimsizing.size3{font-family:KaTeX_Size3}.katex .delimsizing.size4{font-family:KaTeX_Size4}.katex .delimsizing.mult .delim-size1>span{font-family:KaTeX_Size1}.katex .delimsizing.mult .delim-size4>span{font-family:KaTeX_Size4}.katex .nulldelimiter{display:inline-block;width:.12em}.katex .delimcenter,.katex .op-symbol{position:relative}.katex .op-symbol.small-op{font-family:KaTeX_Size1}.katex .op-symbol.large-op{font-family:KaTeX_Size2}.katex .katex-accent>.vlist-t,.katex .op-limits>.vlist-t{text-align:center}.katex .katex-accent .accent-body{position:relative}.katex .katex-accent .accent-body:not(.accent-full){width:0}.katex .katex-overlay{display:block}.katex .mtable .vertical-separator{display:inline-block;min-width:1px}.katex .mtable .arraycolsep{display:inline-block}.katex .mtable .col-align-c>.vlist-t{text-align:center}.katex .mtable .col-align-l>.vlist-t{text-align:left}.katex .mtable .col-align-r>.vlist-t{text-align:right}.katex .svg-align{text-align:left}.katex svg{fill:currentColor;stroke:currentColor;display:block;height:inherit;position:absolute;width:100%}.katex svg path{stroke:none}.katex svg{fill-rule:nonzero;fill-opacity:1;stroke-width:1;stroke-linecap:butt;stroke-linejoin:miter;stroke-miterlimit:4;stroke-dasharray:none;stroke-dashoffset:0;stroke-opacity:1}.katex img{border-style:none;max-height:none;max-width:none;min-height:0;min-width:0}.katex .katex-stretchy{display:block;overflow:hidden;position:relative;width:100%}.katex .katex-stretchy:after,.katex .katex-stretchy:before{content:""}.katex .hide-tail{overflow:hidden;position:relative;width:100%}.katex .halfarrow-left{left:0;overflow:hidden;position:absolute;width:50.2%}.katex .halfarrow-right{overflow:hidden;position:absolute;right:0;width:50.2%}.katex .brace-left{left:0;overflow:hidden;position:absolute;width:25.1%}.katex .brace-center{left:25%;overflow:hidden;position:absolute;width:50%}.katex .brace-right{overflow:hidden;position:absolute;right:0;width:25.1%}.katex .x-arrow-pad{padding:0 .5em}.katex .cd-arrow-pad{padding:0 .55556em 0 .27778em}.katex .mover,.katex .munder,.katex .x-arrow{text-align:center}.katex .boxpad{padding:0 .3em}.katex .fbox,.katex .fcolorbox{border:.04em solid;box-sizing:border-box}.katex .cancel-pad{padding:0 .2em}.katex .cancel-lap{margin-left:-.2em;margin-right:-.2em}.katex .katex-sout{border-bottom-style:solid;border-bottom-width:.08em}.katex .angl{border-right:.049em solid;border-top:.049em solid;box-sizing:border-box;margin-right:.03889em}.katex .anglpad{padding:0 .03889em}.katex .reflectbox{display:inline-block;transform:scaleX(-1)}.katex .eqn-num:before{content:"(" counter(katexEqnNo) ")";counter-increment:katexEqnNo}.katex .mml-eqn-num:before{content:"(" counter(mmlEqnNo) ")";counter-increment:mmlEqnNo}.katex .mtr-glue{width:50%}.katex .cd-vert-arrow{display:inline-block;position:relative}.katex .cd-label-left{display:inline-block;position:absolute;right:calc(50% + .3em);text-align:left}.katex .cd-label-right{display:inline-block;left:calc(50% + .3em);position:absolute;text-align:right}.katex-display{display:block;margin:1em 0;text-align:center}.katex-display>.katex{display:block;text-align:center;white-space:nowrap}.katex-display>.katex>.katex-html{display:block;position:relative}.katex-display>.katex>.katex-html>.katex-tag{position:absolute;right:0}.katex-display.leqno>.katex>.katex-html>.katex-tag{left:0;right:auto}.katex-display.fleqn>.katex{padding-left:2em;text-align:left}body{counter-reset:katexEqnNo mmlEqnNo} | |
| </style> | |
| <script>(function() { | |
| function __mumInit() { | |
| // Restore external links that htmlpreview.github.io's loader rewrote to | |
| // in-page anchors (it treats any href containing '#' as a local anchor). | |
| // Runs now (htmlpreview re-executes this script after its rewrite), again | |
| // after delays, and on click as a last line of defense. | |
| function restoreOrigHrefs() { | |
| document.querySelectorAll('a[data-orig-href]').forEach(function(a) { | |
| var orig = a.getAttribute('data-orig-href'); | |
| if (orig && a.getAttribute('href') !== orig) a.setAttribute('href', orig); | |
| }); | |
| } | |
| restoreOrigHrefs(); | |
| setTimeout(restoreOrigHrefs, 100); | |
| setTimeout(restoreOrigHrefs, 1000); | |
| document.addEventListener('click', function(e) { | |
| var t = e.target && e.target.closest ? e.target.closest('a[data-orig-href]') : null; | |
| if (t) { var o = t.getAttribute('data-orig-href'); if (o && t.getAttribute('href') !== o) t.setAttribute('href', o); } | |
| }, true); | |
| var vt = document.getElementById('viewing-time'); | |
| if (vt) { | |
| function pad(n) { return n < 10 ? '0' + n : n; } | |
| function updateTime() { | |
| var d = new Date(); | |
| vt.textContent = d.getUTCFullYear() + '-' + pad(d.getUTCMonth()+1) + '-' + pad(d.getUTCDate()) + ' ' + pad(d.getUTCHours()) + ':' + pad(d.getUTCMinutes()) + ' UTC'; | |
| } | |
| updateTime(); | |
| setInterval(updateTime, 60000); | |
| } | |
| document.querySelectorAll('a[data-hover-title]').forEach(function(a) { | |
| var card = document.createElement('span'); | |
| card.className = 'hover-card'; | |
| var title = a.getAttribute('data-hover-title') || ''; | |
| var status = a.getAttribute('data-hover-status') || ''; | |
| var owner = a.getAttribute('data-hover-owner') || ''; | |
| var desc = a.getAttribute('data-hover-desc') || ''; | |
| var statusCls = 'hc-status hc-status-' + status.toLowerCase().replace(/[^a-z]/g, ''); | |
| var html = '<div class="hc-title">' + title + '</div>'; | |
| var meta = []; | |
| if (status) meta.push('<span class="' + statusCls + '">' + status + '</span>'); | |
| if (owner) meta.push(owner); | |
| if (meta.length) html += '<div class="hc-meta">' + meta.join(' ') + '</div>'; | |
| if (desc) html += '<div class="hc-desc">' + desc + '</div>'; | |
| card.innerHTML = html; | |
| a.appendChild(card); | |
| }); | |
| document.querySelectorAll('a.cite').forEach(function(a) { | |
| var href = a.getAttribute('href') || ''; | |
| if (!href.startsWith('#ref-')) return; | |
| var li = document.getElementById(href.slice(1)); | |
| if (!li) return; | |
| var body = li.querySelector('.ref-body'); | |
| if (!body) return; | |
| var refUrl = body.querySelector('.ref-url a'); | |
| var refHref = refUrl ? refUrl.getAttribute('href') : ''; | |
| var hoverSource = refHref ? document.querySelector('a[data-hover-title][href="' + refHref + '"]') : null; | |
| if (hoverSource) { | |
| var card = document.createElement('span'); | |
| card.className = 'hover-card'; | |
| var t = hoverSource.getAttribute('data-hover-title') || ''; | |
| var s = hoverSource.getAttribute('data-hover-status') || ''; | |
| var o = hoverSource.getAttribute('data-hover-owner') || ''; | |
| var d = hoverSource.getAttribute('data-hover-desc') || ''; | |
| var sc = 'hc-status hc-status-' + s.toLowerCase().replace(/[^a-z]/g, ''); | |
| var h = '<div class="hc-title">' + t + '</div>'; | |
| var m = []; | |
| if (s) m.push('<span class="' + sc + '">' + s + '</span>'); | |
| if (o) m.push(o); | |
| if (m.length) h += '<div class="hc-meta">' + m.join(' ') + '</div>'; | |
| if (d) h += '<div class="hc-desc">' + d + '</div>'; | |
| card.innerHTML = h; | |
| a.appendChild(card); | |
| } else { | |
| var title = body.querySelector('.ref-title'); | |
| if (!title) return; | |
| var card = document.createElement('span'); | |
| card.className = 'cite-card'; | |
| card.textContent = title.textContent; | |
| a.appendChild(card); | |
| } | |
| }); | |
| } | |
| if (document.readyState !== 'loading') { __mumInit(); } | |
| else { document.addEventListener('DOMContentLoaded', __mumInit); } | |
| })();</script> | |
| </head> | |
| <body> | |
| <div class="page"> | |
| <header class="paper-header"> | |
| <h1>IFM vs. Marin: pretraining, data and post-training stacks</h1> | |
| <div class="paper-meta">Posed by <span class="author">hammer</span>, answered by <a class="mumwelt-link" href="https://github.com/marin-community/mumwelt">mumwelt</a> · <span class="date-label" title="Generated 2026-09-29 08:01 UTC">Published 2026-09-29</span></div> | |
| </header> | |
| <div class="prompt-box"> | |
| <div class="prompt-label">?</div> | |
| <div class="prompt-text">Do a detailed analysis of https://github.com/ifm-ai/xllm, a new library for specifying and training large LLMs, and compare it to what we have in Levanter. Consider flexibility in specifying new architectures, completeness of the linear and other scalable attention mechanisms implemented, FP8 and FP4 training, completeness of kernels for different GPUs, ability to train on non-GPUs, dimensions of parallelism implemented, ease of evolution by agents, code complexity, test coverage, performance, and any other measure important for a team building frontier LLMs that needs to choose between Levanter and xLLM. Then compare the data curation and mixing code in https://github.com/ifm-ai/pretraining-data-toolkit with Marin's Datakit, and https://github.com/ifm-ai/RL360 with MarinSkyRL and other post-training code from Levanter and Marin, including what IFM's talk at https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf says about data curation and mixing.</div> | |
| </div> | |
| <div class="content"> | |
| <p><em>Code read at IFM's xLLM <a | |
| href="https://github.com/ifm-ai/xllm/tree/889db388d021c92164e7a675f6884dd40ea117c7"><code>889db38</code></a>, | |
| xattn <a | |
| href="https://github.com/ifm-ai/xattn/tree/0d6d73b37acc644d5559007fad370185b78ddbb5"><code>0d6d73b</code></a>, | |
| xbridges <a | |
| href="https://github.com/ifm-ai/xbridges/tree/227f3fd26e30b024ce09bc7c556c64616c13276e"><code>227f3fd</code></a>, | |
| pretraining-data-toolkit <a | |
| href="https://github.com/ifm-ai/pretraining-data-toolkit/tree/cd93e116a8217c9e9a45e104b0ea28ac653a998f"><code>cd93e11</code></a> | |
| and RL360 <a | |
| href="https://github.com/ifm-ai/RL360/tree/5b548b6948e6e78af99307a38ee96057ea571faa"><code>5b548b6</code></a>, | |
| and at Marin <code>main</code> <a | |
| href="https://github.com/marin-community/marin/tree/4aa26ec73d248b974766f8940a896a8c888d8269"><code>4aa26ec</code></a> | |
| and MarinSkyRL <a | |
| href="https://github.com/marin-community/MarinSkyRL/tree/0cdccc9229f47c8eadff90aa2033a5b4604266aa"><code>0cdccc9</code></a>, | |
| all as of 2026-09-28 or 2026-09-29. Marin history comes from its GitHub, | |
| Discord and W&B records and its <a | |
| href="https://openathena.ai/blog/marin-data-pipeline-overview/">Datakit | |
| blog post</a>; IFM history from its blog, model and dataset cards, | |
| public W&B logs and a <a | |
| href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf">talk | |
| on its data work</a>. Nothing was run on a GPU or TPU: every speed below | |
| was reported or logged by its authors.</em></p> | |
| <h2 id="summary">Summary</h2> | |
| <p><a href="https://github.com/ifm-ai/xllm">xLLM</a> is the PyTorch | |
| framework that MBZUAI's Institute of Foundation Models (IFM) used to | |
| pretrain its K2 Horizon models, which range from a 0.9B dense model to a | |
| 375B-A23B mixture-of-experts (MoE) model. IFM opened the repository on | |
| 2026-09-02 and pushed the code on 2026-09-28. <em>Levanter</em> here | |
| means Marin's JAX stack: the Levanter library, the Haliax named-tensor | |
| library, and Grug, the copy-and-edit model code now training Marin's | |
| 535B-A23B MoE model on 704 GB200 GPUs. Sections 13 and 14 extend the | |
| comparison to data curation and post-training: IFM's | |
| pretraining-data-toolkit and RL360 against Marin's Datakit and | |
| MarinSkyRL.</p> | |
| <p>Levanter is the stronger base for frontier pretraining. It has what | |
| xLLM lacks: Muon-family optimizers, expert parallelism across a 64-GPU | |
| NVLink domain, proven use on Blackwell GPUs, and TPU support. Neither | |
| stack's production runs use pipeline parallelism, gradient accumulation | |
| or FP8, but Levanter has working code for all three and xLLM's public | |
| code has none. Levanter also tests far more and is built for agents to | |
| change.</p> | |
| <p>xLLM has trained dense models at scale. IFM's public logs show an | |
| internal version of it pretraining a 7B model on 21.9T tokens at about | |
| 48.6% model FLOPs utilization (MFU) by its own formula, and extending | |
| context to 512K tokens. It supports sparse MoE and probably pretrained | |
| the 375B-A23B as well, though IFM has published no throughput for that | |
| run. Above its kernels it tests only its data pipeline, it has no CI, | |
| and its most original component, a long-context layer called Gekko, | |
| appears in none of the released models.</p> | |
| <p>Levanter has weaknesses too. Its GPU speed depends on a patched XLA | |
| plugin and exactly pinned, fast-moving kernel packages, and the hero's | |
| GB200 path rests mainly on one engineer. No GPU tests run on pull | |
| requests. And Marin has no 7–8B dense pretraining measurement on GPUs to | |
| set beside xLLM's numbers.</p> | |
| <p>For data curation, Marin has the more complete code and IFM has | |
| released far more data. IFM's public data toolkit covers the middle of | |
| its pipeline: labeling documents with public classifiers, sorting them | |
| into quality categories and shuffling them, in about 3,100 lines with no | |
| tests. The Common Crawl pipeline behind K2 Horizon's web text is public | |
| from IFM's earlier TxT360 work, but the processing of its newer sources, | |
| most of its synthesis code and its mixture search are not. IFM has | |
| released about 13.5 trillion tokens of training data, mostly synthetic. | |
| Marin's Datakit runs from download through deduplication, | |
| decontamination, topic and quality bucketing and tokenization, with | |
| hundreds of tests. Marin rehosts none of its pretraining corpus, several | |
| of whose sources forbid redistribution, and the code that chose its | |
| hero's buckets and mixture weights is not on <code>main</code>. Both | |
| labs found that searching for mixture weights gains less than adding new | |
| or more diverse data.</p> | |
| <p>For post-training, Marin has the more complete RL code and IFM the | |
| results. Both run RL in PyTorch on Megatron and run agentic tasks | |
| through Harbor. IFM's RL360 is a snapshot of five forked projects with | |
| no tests of its own and one demonstration recipe that IFM says has not | |
| been validated end to end; IFM trained its shipped experts with internal | |
| versions of it and other harnesses, and its merge code is private. | |
| MarinSkyRL, Marin's fork of SkyRL, has about 2,300 tests but changes | |
| weekly: in the four days before the commit read here, it deleted its | |
| FSDP trainer and replaced its training loop. Because Marin pretrains and | |
| fine-tunes in JAX, its RL needs a PyTorch port of each model, a cost IFM | |
| avoids. IFM has shipped post-trained models at six sizes; Marin's first | |
| official post-trained model has been pushed back a week from its planned | |
| October 6 debut.</p> | |
| <p><strong>Recommendation.</strong> Marin should stay on Levanter, run | |
| one same-hardware dense benchmark against xLLM, count causal attention | |
| in its MFU, and study IFM's long-context schedule, Gekko and MoVA. From | |
| the data and post-training comparison, it should test IFM's released | |
| datasets, favor new data over further mixture search, move its hero's | |
| data and SFT code onto <code>main</code>, and compare weight merging | |
| with distillation. For other teams, the hardware decides the pretraining | |
| stack: on TPUs or GB200 racks, choose Levanter; on Hopper clusters under | |
| Slurm, for dense models or long-context research, xLLM is a credible | |
| start; for frontier MoE on Hopper, xLLM has probably trained a 375B-A23B | |
| model but published no efficiency data, so benchmark it against | |
| TorchTitan and Megatron-Core. For data, Datakit is the more complete | |
| pipeline, and IFM's datasets are worth more than its toolkit. For RL, | |
| start from upstream Miles or SkyRL rather than either lab's fork.</p> | |
| <table> | |
| <thead> | |
| <tr> | |
| <th>Dimension</th> | |
| <th>IFM (xLLM, data toolkit, RL360)</th> | |
| <th>Marin (Levanter + Grug, Datakit, MarinSkyRL)</th> | |
| <th>Edge</th> | |
| </tr> | |
| </thead> | |
| <tbody> | |
| <tr> | |
| <td>New architectures</td> | |
| <td>Two model types chosen by flags; new blocks train through autograd, | |
| but the fast path needs a hand-written backward</td> | |
| <td>Copy a Grug template and edit plain JAX; autodiff supplies | |
| gradients</td> | |
| <td>Levanter</td> | |
| </tr> | |
| <tr> | |
| <td>Optimizers</td> | |
| <td>AdamW only</td> | |
| <td>16, including Muon and MuonH (used by the hero)</td> | |
| <td>Levanter</td> | |
| </tr> | |
| <tr> | |
| <td>Attention and linear mixers</td> | |
| <td>Full attention in every released model; Gekko hybrid unused; no | |
| sliding window</td> | |
| <td>Sliding-window/global hybrid in production; Gated DeltaNet and | |
| Mamba-3 unwired; KDA on branches</td> | |
| <td>Neither is broad</td> | |
| </tr> | |
| <tr> | |
| <td>Long-context training</td> | |
| <td>3.7B and 7B models mid-trained to 512K tokens</td> | |
| <td>67B MoE continued at 262K on TPUs; 262K on GPUs only in tests</td> | |
| <td>xLLM</td> | |
| </tr> | |
| <tr> | |
| <td>FP8 / FP4</td> | |
| <td>None</td> | |
| <td>FP8 for classic dense layers; MoE FP8 tried and shelved; no FP4</td> | |
| <td>Neither</td> | |
| </tr> | |
| <tr> | |
| <td>GPU kernels</td> | |
| <td>Set up for Hopper; no Blackwell run reported</td> | |
| <td>Tuned for Blackwell; older Hopper attention; patched XLA plugin</td> | |
| <td>Levanter on Blackwell; likely xLLM on Hopper</td> | |
| </tr> | |
| <tr> | |
| <td>Non-NVIDIA hardware</td> | |
| <td>None</td> | |
| <td>TPU v4 to v6e</td> | |
| <td>Levanter</td> | |
| </tr> | |
| <tr> | |
| <td>Parallelism</td> | |
| <td>FSDP, TP ≤ 8, CP; EP tied to TP; no PP</td> | |
| <td>FSDP, EP64 in production, CP, cross-rack DP; PP in benchmarks</td> | |
| <td>Levanter</td> | |
| </tr> | |
| <tr> | |
| <td>Sparse MoE</td> | |
| <td>Top-k routing, shared experts, no token dropping; EP within one | |
| node; probably trained a 375B-A23B, efficiency unpublished</td> | |
| <td>Eight transports, capacity controls, EP64; 535B-A23B hero at 26.68% | |
| MFU</td> | |
| <td>Levanter</td> | |
| </tr> | |
| <tr> | |
| <td>Operations</td> | |
| <td>Slurm or torchrun, a shared filesystem; resume only on the same | |
| layout</td> | |
| <td>Kubernetes and TPU pools; object-store checkpoints that restore onto | |
| any mesh</td> | |
| <td>Levanter</td> | |
| </tr> | |
| <tr> | |
| <td>Performance</td> | |
| <td>About 48.6% MFU, derived from logs, for a dense 7B over 21.9T | |
| tokens</td> | |
| <td>26.68% MFU for a 535B-A23B MoE on 704 GB200s</td> | |
| <td>No like-for-like data</td> | |
| </tr> | |
| <tr> | |
| <td>Evolution by agents</td> | |
| <td>Small, but no CI or agent docs; no model module imports without a | |
| CUDA build</td> | |
| <td>Agent instructions, 40 skills and CPU-runnable tests; agents already | |
| change the stack</td> | |
| <td>Levanter</td> | |
| </tr> | |
| <tr> | |
| <td>Code size</td> | |
| <td>21.5K lines Python, 14.8K own C++/CUDA, plus 26.2K in xattn</td> | |
| <td>84K lines Python, 1.3K CUDA, plus pinned external kernels</td> | |
| <td>Mixed</td> | |
| </tr> | |
| <tr> | |
| <td>Tests</td> | |
| <td>62 test functions; kernels and data pipeline only; no CI</td> | |
| <td>About 1,900 test functions; CPU and TPU tests on pull requests; | |
| daily GPU and TPU canaries, but no GPU tests on pull requests</td> | |
| <td>Levanter</td> | |
| </tr> | |
| <tr> | |
| <td>Data curation</td> | |
| <td>Toolkit labels, sorts and shuffles (3.1K lines, no tests); newer | |
| sources' processing unpublished; about 13.5T tokens of data | |
| released</td> | |
| <td>Datakit covers download to tokenized store, including deduplication | |
| and decontamination (360 tests); corpus not rehosted; code that chose | |
| the hero's buckets off <code>main</code></td> | |
| <td>Marin for code, IFM for released data</td> | |
| </tr> | |
| <tr> | |
| <td>Data mixing</td> | |
| <td>Weights passed to the trainer, none published; search method | |
| described only in a talk, which reports small gains for 375K | |
| GPU-hours</td> | |
| <td>Proxy-model swarms over 200 topic × quality buckets; latest re-mix | |
| 1.20× over the previous mix on a ladder; weights published, search code | |
| off <code>main</code></td> | |
| <td>Marin</td> | |
| </tr> | |
| <tr> | |
| <td>Post-training</td> | |
| <td>SFT in xLLM; RL360 snapshot of Miles, Megatron-LM, SGLang, SMG and | |
| Harbor with one demo recipe and no tests of its own; merge code private; | |
| six post-trained models shipped, four with RL</td> | |
| <td>SFT in JAX, partly on unmerged branches; MarinSkyRL (SkyRL fork, | |
| Megatron only) with about 2,300 tests and nightly GPU runs; first | |
| release pending</td> | |
| <td>Marin for RL code, IFM for results</td> | |
| </tr> | |
| </tbody> | |
| </table> | |
| <p><em>Terms.</em> MFU: the share of peak hardware FLOPs spent on the | |
| model's own arithmetic. FSDP, DP, TP, CP, EP, PP: fully sharded data, | |
| data, tensor, context, expert and pipeline parallelism; EP64 means | |
| 64-way. Hopper and Blackwell: NVIDIA's H100/H200 (SM90) and B200/GB200 | |
| (SM100) generations. XLA: the compiler behind JAX. GQA: grouped-query | |
| attention. KDA: Kimi Delta Attention, a linear-attention layer. SFT: | |
| supervised fine-tuning. RL: reinforcement learning; RLVR uses verifiable | |
| rewards, such as passing tests. A rollout is one sampled answer or agent | |
| episode. GRPO and RLOO: RL methods that score each rollout against | |
| others for the same prompt; DAPO adds asymmetric clipping and drops | |
| uninformative prompts. TIS: truncated importance sampling, which | |
| corrects for mismatch between the rollout and training engines. Router | |
| replay: reusing the inference engine's expert choices during training. | |
| On-policy distillation: training a model on its own samples, scored by a | |
| teacher. pass@k: the chance that at least one of k samples is correct. | |
| ISO and RAM: two weight-merging methods named in IFM's model cards. BPB: | |
| bits per byte, a loss that does not depend on the tokenizer. MinHash and | |
| LSH: hashing methods for finding near-duplicate documents. Miles, slime | |
| and SkyRL are RL frameworks, SGLang and vLLM are inference engines, and | |
| Harbor runs agent tasks in sandboxes.</p> | |
| <h2 id="the-two-stacks">The two stacks</h2> | |
| <p><strong>xLLM and K2 Horizon.</strong> IFM released six K2 Horizon | |
| models on 2026-09-03: dense 0.9B, 3.7B, 7B and 32B models, MoVA-36B-A4B | |
| (an MoE that also routes its attention value projections to experts) and | |
| a 375B-A23B MoE, with contexts up to 512K tokens (<a | |
| href="https://huggingface.co/IFM/K2-Horizon-7B">model card</a>). The | |
| launch post calls xLLM <a | |
| href="https://web.archive.org/web/20260913112353/https://ifm.ai/blog/k2/">"our | |
| production-tested training infrastructure"</a>. IFM's public W&B | |
| logs show an internal version of xLLM training the three smallest | |
| models: the runs use xLLM's metric names and also log metrics the public | |
| code cannot produce (<a href="https://wandb.ai/llm360/K2-Horizon-7B">7B | |
| logs</a>). The cards credit xLLM for the larger three as well, and a | |
| 375B run name mentions 256 nodes and 8-way expert parallelism, but no | |
| public log confirms the trainer (section 7 weighs the evidence). | |
| Reinforcement learning ran on separate harnesses, including internal | |
| versions of <a href="https://github.com/ifm-ai/RL360">RL360</a>, a | |
| Megatron-LM-based stack that IFM has released as a snapshot (section | |
| 14). IFM does not disclose K2 Horizon's hardware; logged memory peaks of | |
| about 140 GiB per GPU point to H200s, which IFM used for its previous | |
| model (<a href="https://arxiv.org/abs/2512.06201">K2-V2</a>).</p> | |
| <p>The public history is two commits; the second adds 306 files and 64K | |
| lines in one squash (<a | |
| href="https://github.com/ifm-ai/xllm/commit/889db388d021c92164e7a675f6884dd40ea117c7">commit</a>). | |
| GitHub lists one contributor, Xuezhe (Max) Ma, lead author of the MEGA, | |
| Megalodon and <a href="https://arxiv.org/abs/2601.06463">Gecko</a> | |
| architectures; xLLM's code descends from Ma's Megalodon repository. Two | |
| companion repositories matter: <a | |
| href="https://github.com/ifm-ai/xattn">xattn</a>, attention kernels | |
| derived from FlashAttention-3 (<a | |
| data-orig-href="https://github.com/ifm-ai/xattn/blob/0d6d73b37acc644d5559007fad370185b78ddbb5/THIRD_PARTY_NOTICES#L1-L16" href="https://github.com/ifm-ai/xattn/blob/0d6d73b37acc644d5559007fad370185b78ddbb5/THIRD_PARTY_NOTICES#L1-L16">notices</a>), | |
| and <a href="https://github.com/ifm-ai/xbridges">xbridges</a>, which | |
| converts checkpoints for Hugging Face and vLLM. xLLM needs PyTorch 2.11 | |
| or later and CUDA 12.8 or later, and launches under torchrun or Slurm | |
| (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/README.md#L39-L40" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/README.md#L39-L40">README</a>).</p> | |
| <p><strong>Levanter.</strong> Two model styles share one trainer, data | |
| loader and checkpointer. <em>Classic</em> Levanter builds models from | |
| Haliax named axes, maps them to devices through configuration, and | |
| imports and exports about 15 Hugging Face architectures (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/models/lm_model.py#L154-L195" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/models/lm_model.py#L154-L195">registry</a>). | |
| <em>Grug</em>, added in January 2026, writes each model in plain JAX | |
| with explicit sharding and asks developers to copy a template and edit | |
| it (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/README.md#L1-L27" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/README.md#L1-L27">README</a>). | |
| David Hall gave the reason: <a | |
| href="https://discord.com/channels/1354881461060243556/1462884917292699669/1464708102887706725">"the | |
| motivation is that agents will know jax way better than haliax"</a>. | |
| Grug trains Marin's frontier runs: the 535B-A23B <em>hero</em> (704 | |
| GB200s, 7.96T of a planned 18T tokens by 2026-09-28) and Snowball | |
| 67B-A2B (10T tokens on a TPU v4-2048, <a | |
| data-orig-href="https://github.com/marin-community/marin/issues/6044#issuecomment-5401358100" href="https://github.com/marin-community/marin/issues/6044#issuecomment-5401358100">#6044</a>). | |
| Marin curates pretraining data with Datakit and runs RL with MarinSkyRL, | |
| a PyTorch fork of SkyRL (sections 13 and 14).</p> | |
| <h2 id="1-specifying-new-architectures">1. Specifying new | |
| architectures</h2> | |
| <p><strong>xLLM</strong> picks one of two architectures, | |
| <code>transformer</code> or <code>gekko</code>, from a hard-coded | |
| dictionary (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/build.py#L29-L33" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/build.py#L29-L33">build.py</a>). | |
| Each setting is a single command-line flag; the parser rejects | |
| dictionaries and has no syntax for per-layer lists (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/configuration/configuration.py#L140-L190" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/configuration/configuration.py#L140-L190">configuration.py</a>). | |
| So layers cannot vary, except that the first N may be dense before the | |
| MoE layers begin (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L607-L612" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L607-L612"><code>num_dense_layers</code></a>). | |
| There is no interleaved MoE, no per-layer attention type, and no mixing | |
| of Transformer and Gekko layers.</p> | |
| <p>Each block exists twice: an eager <code>nn.Module</code> that trains | |
| through autograd and serves evaluation, and an optional fused | |
| <code>autograd.Function</code> with a hand-written backward. At equal | |
| settings the eager path is nearly as fast (43.3% against 44.1% MFU on | |
| Llama3-8B), but the fastest configuration, 52.1%, needs the fused path | |
| and its fine-grained recomputation switches. The fused functions take 44 | |
| to 65 positional inputs, and their backward passes run 240 to 578 lines | |
| (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L120-L167" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L120-L167">dense | |
| block</a>). No test compares the two paths, and they already disagree: | |
| the fused path always uses plain additive residuals and FlashAttention, | |
| whatever the configuration asks for (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/fused_blocks/transformer/transformer.py#L117" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/fused_blocks/transformer/transformer.py#L117">residual</a>, | |
| <a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/fused_blocks/mha.py#L73-L75" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/fused_blocks/mha.py#L73-L75">attention</a>). | |
| A block that must also be exported needs separate Hugging Face and vLLM | |
| ports in xbridges. MoVA came to about 2,100 lines across seven files in | |
| two repositories.</p> | |
| <p>xLLM trains with AdamW alone, in one parameter group, so weight decay | |
| also applies to embeddings and norm gains (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/optim/optimizer.py#L24-L32" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/optim/optimizer.py#L24-L32">optimizer.py</a>). | |
| There is no Muon, gradient accumulation, weight averaging, multi-token | |
| prediction or logit z-loss; a router z-loss setting exists, but nothing | |
| reads it (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L352" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L352">config.py</a>).</p> | |
| <p><strong>Levanter.</strong> A Grug variant starts as a copy of a | |
| 275-line model file (<a | |
| href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/base/model.py">base/model.py</a>), | |
| and JAX differentiates it, so each block exists once. In unrolled | |
| variants, per-layer choices are ordinary Python. The hero scans one | |
| compiled block over its 48 layers and passes each layer's choice of | |
| local or global attention in as data (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/moe_hero_ep/model.py#L1298-L1339" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/moe_hero_ep/model.py#L1298-L1339">model.py</a>). | |
| A contract test traces one training step of each variant, except the | |
| pipeline one, using abstract shapes on a CPU mesh (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/tests/test_grug_variant_contracts.py#L248-L288" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/tests/test_grug_variant_contracts.py#L248-L288">test</a>), | |
| and CI posts each new variant's diff against its closest relative. | |
| Sixteen optimizers are registered, including Muon, MuonH, SOAP and Kron | |
| (<a | |
| href="https://github.com/marin-community/marin/tree/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/optim">optim/</a>), | |
| and the trainer supports z-loss and weight averaging.</p> | |
| <p>Grug pays for this freedom in five ways. The hero's model file | |
| repeats about 1,000 lines of its FSDP sibling, and one optimizer file | |
| exists in five copies. Sharding lives in model code. The scanned hero | |
| needs identical parameter shapes in every layer, so it does not support | |
| dense layers before MoE layers. Grug cannot import Hugging Face weights. | |
| And the classic Muon optimizers find matrices by looking for Haliax | |
| <code>Linear</code> layers (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/optim/muon.py#L101-L121" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/optim/muon.py#L101-L121">muon.py</a>), | |
| which Grug models lack; on a Grug model they would silently apply | |
| AdamW-style updates, so Grug needs its own variant (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/optim/grugmuon.py#L4-L8" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/optim/grugmuon.py#L4-L8">grugmuon.py</a>).</p> | |
| <p><em>Verdict: Levanter. In Grug a new block is one Python edit. In | |
| xLLM a new block trains easily, but reaches full speed only through a | |
| second, hand-differentiated implementation that nothing checks.</em></p> | |
| <h2 id="2-attention-linear-mixers-and-long-context">2. Attention, linear | |
| mixers and long context</h2> | |
| <table> | |
| <thead> | |
| <tr> | |
| <th>Mechanism</th> | |
| <th>xLLM</th> | |
| <th>Levanter + Grug</th> | |
| </tr> | |
| </thead> | |
| <tbody> | |
| <tr> | |
| <td>Causal softmax attention, GQA, packed documents</td> | |
| <td>FlashAttention 2, 3 or 4 (external, unpinned)</td> | |
| <td>Splash (TPU); FlashAttention-4 via CuTe (GPU), skipping masked | |
| document blocks</td> | |
| </tr> | |
| <tr> | |
| <td>Sliding window</td> | |
| <td>Absent: window hard-coded to −1 (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/fused_ops/attention/flash.py#L119-L131" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/fused_ops/attention/flash.py#L119-L131">flash.py</a>)</td> | |
| <td>In production: a 2,048-token window on three of every four hero | |
| layers</td> | |
| </tr> | |
| <tr> | |
| <td>Attention sinks, logit soft-capping</td> | |
| <td>Absent</td> | |
| <td>Classic reference and Splash paths only</td> | |
| </tr> | |
| <tr> | |
| <td>Multi-head latent attention (MLA)</td> | |
| <td>Absent</td> | |
| <td>Classic layer, unwired; rejected for the hero (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/6522#issuecomment-4868376304" href="https://github.com/marin-community/marin/issues/6522#issuecomment-4868376304">#6522</a>)</td> | |
| </tr> | |
| <tr> | |
| <td>Sliding-chunk attention plus delta-rule memory (Gekko)</td> | |
| <td>Wired, IFM's own design, used by no released model</td> | |
| <td>Absent</td> | |
| </tr> | |
| <tr> | |
| <td>Gated DeltaNet, KDA</td> | |
| <td>Absent</td> | |
| <td>Gated DeltaNet in pure JAX, unwired; KDA on branches with an | |
| H100-only kernel</td> | |
| </tr> | |
| <tr> | |
| <td>Mamba-style state-space layers</td> | |
| <td>Absent</td> | |
| <td>Mamba-3 and SSD reference code, unwired</td> | |
| </tr> | |
| <tr> | |
| <td>Published Gecko layer (complex moving average, Megalodon line)</td> | |
| <td>Present, broken, unwired</td> | |
| <td>Absent</td> | |
| </tr> | |
| <tr> | |
| <td>Short causal convolution</td> | |
| <td>In Gekko (Triton)</td> | |
| <td>In the hero (Triton on GPU, XLA on TPU)</td> | |
| </tr> | |
| <tr> | |
| <td>Context parallelism</td> | |
| <td>Gathers all keys and values; Gekko passes state rank to rank; no | |
| public example or test</td> | |
| <td>Gathers all keys and values; tested at 262K tokens on 64 GB200s</td> | |
| </tr> | |
| <tr> | |
| <td>RoPE scaling</td> | |
| <td>Absent</td> | |
| <td>Classic: YaRN and Llama 3 scaling; Grug: a temperature on | |
| queries</td> | |
| </tr> | |
| </tbody> | |
| </table> | |
| <p>Every released K2 Horizon model is a plain Transformer with full | |
| causal attention, and every example trains at 8,192 tokens without | |
| context parallelism (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/k2-horizon-7B.sh#L23-L57" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/k2-horizon-7B.sh#L23-L57">example</a>). | |
| For K2 Horizon, the README's <a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/README.md#L5" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/README.md#L5">"extra-long | |
| contexts"</a> means full attention at long sequence lengths, and IFM has | |
| used it at scale. The 7B was mid-trained on 1.1T tokens at 32K, 498B at | |
| 128K and 309B at 512K, then fine-tuned at 512K. Across its four 512K | |
| phases it ran at medians of 632 to 991 tokens/s/GPU, 33–51% of peak by | |
| xLLM's formula, which credits attention work that document masking skips | |
| (<a href="https://wandb.ai/llm360/K2-Horizon-7B">7B logs</a>). The logs | |
| do not say how each sequence was split across GPUs, and IFM has not | |
| released the mid-training data.</p> | |
| <p>Marin's longest-context training ran at 262K tokens on TPUs: about | |
| 67B tokens of continued pretraining for the 67B-A2B, then fine-tuning | |
| (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/6044#issuecomment-5401431533" href="https://github.com/marin-community/marin/issues/6044#issuecomment-5401431533">#6044</a>, | |
| <a | |
| href="https://github.com/marin-community/marin/issues/8954">#8954</a>). | |
| On GPUs, Marin's 262K runs are tests at about 10% MFU (<a | |
| href="https://github.com/marin-community/marin/pull/9119">#9119</a>).</p> | |
| <p>Gekko is xLLM's most original component. Each layer normalizes with | |
| decaying running statistics and applies short convolutions to queries, | |
| keys and values. It then adds two branches: softmax attention over the | |
| current and previous chunk, and an <em>adaptive working memory</em> that | |
| summarizes older chunks with a delta-rule update (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/gated_delta_attention.py#L326-L451" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/gated_delta_attention.py#L326-L451">gated_delta_attention.py</a>). | |
| It adapts Ma's published <a | |
| href="https://arxiv.org/abs/2601.06463">Gecko</a> architecture, which | |
| was pretrained at 7B on 2T tokens: it keeps Gecko's decaying norm, chunk | |
| attention and memory update, drops its complex moving average, and adds | |
| short convolutions. Under context parallelism each rank hands its memory | |
| to the next. Gekko's parts have reference tests, but no test covers a | |
| whole block and no released model uses it. Its memory update runs as a | |
| Python loop over chunks (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/fused_ops/adaptive_working_memory.py#L99-L241" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/fused_ops/adaptive_working_memory.py#L99-L241">adaptive_working_memory.py</a>), | |
| and because xLLM always sends a document mask, xattn's fast | |
| chunk-attention kernel runs only on Hopper.</p> | |
| <p>Three attention paths do not run. The <code>xattn</code> backend | |
| raises <code>NotImplementedError</code> when a Transformer is built (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/causal_attention.py#L53-L54" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/causal_attention.py#L53-L54">causal_attention.py</a>). | |
| The <code>swift</code> backend cannot unpack the segment data that | |
| packed documents produce (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L719-L725" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L719-L725">producer</a>, | |
| <a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/causal_attention.py#L124-L125" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/causal_attention.py#L124-L125">consumer</a>) | |
| and stops above 8,192 keys (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/csrc/softmax.cuh#L387-L391" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/csrc/softmax.cuh#L387-L391">softmax.cuh</a>). | |
| And the module for the published Gecko layer calls functions with | |
| outdated signatures (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moving_average_gated_attention.py#L199-L208" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moving_average_gated_attention.py#L199-L208">moving_average_gated_attention.py</a>).</p> | |
| <p>Levanter's production design is softmax attention tuned by ablation: | |
| sliding windows with a global layer every fourth layer that drops | |
| positional encoding, query-key norms, per-head gates and short | |
| convolutions. The same ablation found full multi-head attention better | |
| than the grouped-query heads the hero kept (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/8227#issuecomment-5310665092" href="https://github.com/marin-community/marin/issues/8227#issuecomment-5310665092">#8227</a>). | |
| Its linear-attention work is younger. On 8×H100 ladder models with up to | |
| 291M active parameters, KDA in the local layers gave 1.13–1.26× the | |
| baseline's compute efficiency, and a hybrid adding MLA global layers and | |
| attention residuals gave 1.14–1.43× (<a | |
| href="https://github.com/marin-community/marin/issues/9438">#9438</a>, | |
| <a | |
| href="https://github.com/marin-community/marin/issues/9451">#9451</a>). | |
| That work lives on branches, with an H100 kernel and none for TPU; TPU | |
| kernels for Gated DeltaNet are in draft (<a | |
| href="https://github.com/marin-community/marin/pull/9350">#9350</a>). | |
| One trap: the hero passes its window only through FlashAttention-4's | |
| bounds, so if its model runs on the reference or Splash attention path, | |
| the sliding-window layers silently become full attention (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/moe_hero_ep/model.py#L1283-L1329" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/moe_hero_ep/model.py#L1283-L1329">model.py</a>).</p> | |
| <p>PyTorch offers an advantage xLLM leaves unused. The Flash Linear | |
| Attention library has tuned kernels for dozens of linear-attention | |
| variants, and xLLM imports it only for a causal convolution. Marin had | |
| to write its own chunked KDA kernel for JAX in September (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/9438#issuecomment-5826696333" href="https://github.com/marin-community/marin/issues/9438#issuecomment-5826696333">#9438</a>).</p> | |
| <p><em>Verdict: neither stack offers a broad, production-grade set of | |
| scalable mixers. Levanter runs the richer attention design in production | |
| and tests more candidates; xLLM has one complete hybrid that no released | |
| model uses. On long context, xLLM has done more in production, on | |
| smaller models.</em></p> | |
| <h2 id="3-fp8-and-fp4-training">3. FP8 and FP4 training</h2> | |
| <p><strong>xLLM has no low-precision training path.</strong> The config | |
| accepts only <code>bf16</code>, <code>fp16</code> or <code>fp32</code> | |
| (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L414" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L414">config.py</a>). | |
| xLLM copied part of NVIDIA's Transformer Engine <a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/csrc/transformer_engine/NOTICE#L6-L9" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/csrc/transformer_engine/NOTICE#L6-L9">"for | |
| its cuBLAS grouped matrix multiplication backend"</a>; the FP8 code in | |
| that copy never runs, because xLLM passes no scaling factors (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/csrc/transformer_engine/common.cpp#L37-L61" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/csrc/transformer_engine/common.cpp#L37-L61">common.cpp</a>). | |
| FP4 exists only as type definitions.</p> | |
| <p><strong>Levanter has FP8 code that no production run uses.</strong> | |
| Haliax provides an FP8 matmul with E4M3 forward values, E5M2 gradients | |
| and per-tensor delayed scaling, wired only for classic dense | |
| <code>Linear</code> layers (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/haliax/src/haliax/quantization.py#L180-L270" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/haliax/src/haliax/quantization.py#L180-L270">quantization.py</a>). | |
| The MoE layer and all of Grug lack it. Marin measured MoE FP8 carefully, | |
| found modest gains, and parked it:</p> | |
| <ul> | |
| <li>18B MoE, 8 H100s per arm: FP8 raised MFU from 15.1% to 15.9%, with | |
| final loss within 0.004 (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/7298#issuecomment-5010209991" href="https://github.com/marin-community/marin/issues/7298#issuecomment-5010209991">#7298</a>).</li> | |
| <li>64 GB200s, single runs: MXFP8 was 1.31× faster than BF16 with 8-way | |
| expert parallelism (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/7282#issuecomment-5017713906" href="https://github.com/marin-community/marin/issues/7282#issuecomment-5017713906">#7282</a>) | |
| but 0.75× as fast with FSDP alone (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/7282#issuecomment-5037422197" href="https://github.com/marin-community/marin/issues/7282#issuecomment-5037422197">#7282</a>).</li> | |
| <li>32 GB200s, 66B tokens: MXFP8 gave 7.2% more throughput, but final | |
| eval loss was 0.056% worse and BF16 won all 32 paired evaluations (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/7271#issuecomment-5036746436" href="https://github.com/marin-community/marin/issues/7271#issuecomment-5036746436">#7271</a>).</li> | |
| <li>The owner closed the MoE FP8 pull requests in August: <a | |
| data-orig-href="https://github.com/marin-community/marin/pull/6880#issuecomment-5402446303" href="https://github.com/marin-community/marin/pull/6880#issuecomment-5402446303">"Closing | |
| since FP8 is not currently a priority."</a> A September retry on 8 H100s | |
| was parked at a best gain of about 6% (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/9438#issuecomment-5827536865" href="https://github.com/marin-community/marin/issues/9438#issuecomment-5827536865">#9438</a>), | |
| though on 2026-09-18 David Hall wrote of FP8 expert matmuls, <a | |
| data-orig-href="https://github.com/marin-community/marin/issues/6699#issuecomment-5727082751" href="https://github.com/marin-community/marin/issues/6699#issuecomment-5727082751">"still | |
| think we should do this"</a>.</li> | |
| </ul> | |
| <p>On FP4, David Hall wrote: <a | |
| data-orig-href="https://github.com/marin-community/marin/issues/7403#issuecomment-5050343888" href="https://github.com/marin-community/marin/issues/7403#issuecomment-5050343888">"I | |
| strongly think we should stay away from nvfp4 unless we have very very | |
| good reason to take on the risk"</a>. The hero computes in BF16 with | |
| FP32 weights.</p> | |
| <p><em>Verdict: neither trains in FP8 or FP4. Levanter has a working FP8 | |
| matmul and a record of measurements. xLLM would need new work: PyTorch's | |
| FP8 libraries could cover its eager modules, but each fused block would | |
| need hand changes.</em></p> | |
| <h2 id="4-kernels-for-different-gpus">4. Kernels for different GPUs</h2> | |
| <p><strong>xLLM</strong> leaves attention to FlashAttention: FA2 by | |
| default, FA3 or FA4 if an environment variable is set before import (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/fused_ops/attention/flash.py#L7-L23" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/fused_ops/attention/flash.py#L7-L23">flash.py</a>). | |
| Every example sets FA3, which runs only on Hopper. Its own 14.8K lines | |
| of C++/CUDA mostly implement norms, moving-average scans and FFT | |
| convolutions for the Megalodon and Gekko line, and contain no | |
| Hopper-specific code; the hand-written kernels leave matrix products to | |
| cuBLAS. On the released Transformer models they contribute only grouped | |
| RMSNorm, plus an optional grouped GEMM (one cuBLASLt call per expert) | |
| and a Triton kernel for MoE output. About half the registered ops are | |
| never reached. The Dockerfile compiles for SM90 alone (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/Dockerfile#L22" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/Dockerfile#L22">Dockerfile</a>), | |
| and xattn builds for SM80 through SM90 and says <a | |
| data-orig-href="https://github.com/ifm-ai/xattn/blob/0d6d73b37acc644d5559007fad370185b78ddbb5/README.md#L84-L88" href="https://github.com/ifm-ai/xattn/blob/0d6d73b37acc644d5559007fad370185b78ddbb5/README.md#L84-L88">"SM100 | |
| and newer architectures are currently unsupported"</a>. On Blackwell, | |
| xLLM would depend on FA4 alone; its attention test can select FA4, but | |
| IFM reports no Blackwell run. It uses neither <code>torch.compile</code> | |
| nor CUDA graphs.</p> | |
| <p>xLLM's dependency risks are quieter than Levanter's but real: | |
| FlashAttention and Flash Linear Attention are unpinned, the Transformer | |
| Engine copy records no upstream version, its FSDP1 code imports private | |
| PyTorch internals (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/distributed/fully_sharded_data_parallel.py#L10-L13" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/distributed/fully_sharded_data_parallel.py#L10-L13">fully_sharded_data_parallel.py</a>), | |
| and xbridges' own vLLM path requires exactly vLLM 0.24.0.</p> | |
| <p><strong>Levanter</strong> tunes for Blackwell. The hero uses upstream | |
| FlashAttention-4 kernels for SM100 (<a | |
| href="https://github.com/marin-community/marin/pull/9332">#9332</a>), | |
| QuACK kernels for the expert matrix multiplies and for Muon, and Triton | |
| kernels for token gathering and short convolutions. Its expert | |
| all-to-all needs a patched XLA plugin built only for GB200's ARM hosts | |
| (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/pyproject.toml#L189-L192" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/pyproject.toml#L189-L192">pin</a>). | |
| On Hopper, the attention forward pass is Marin's port of CUTLASS's | |
| Ampere FlashAttention-2 example, without Hopper's asynchronous | |
| tensor-core instructions (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/grug/attention/_fa4_cute_kernels.py#L32-L60" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/grug/attention/_fa4_cute_kernels.py#L32-L60">_fa4_cute_kernels.py</a>); | |
| I found no measurement of it against FA3. TPUs use upstream Splash | |
| attention and Megablox grouped matmul.</p> | |
| <p>The price is maintenance. Marin pins fast-moving kernel packages | |
| exactly (a FlashAttention-4 beta, QuACK 0.6.4, CUTLASS DSL 4.6.2), | |
| rebuilds an XLA fork when JAX changes, and logged about a dozen kernel | |
| and runtime incidents, from corruptions to hangs, between July and | |
| September. Two examples: a deadlock inside a QuACK grouped GEMM that | |
| intermittently hung the hero (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/8870#issuecomment-5687716531" href="https://github.com/marin-community/marin/issues/8870#issuecomment-5687716531">#8870</a>), | |
| and a Triton grouped-matmul bug that gave wrong rows on GB200 (69% of | |
| rows in a test layer) and stayed open for five weeks (<a | |
| href="https://github.com/marin-community/marin/issues/7484">#7484</a>, | |
| <a href="https://github.com/marin-community/marin/pull/8610">#8610</a>). | |
| A few engineers own the kernels, and the hero's GB200 path rests mainly | |
| on one of them, working with an agent. IFM's incident history is | |
| private, so the two records cannot be compared; its 7B run restarted 95 | |
| times.</p> | |
| <p><em>Verdict: Levanter on Blackwell. On Hopper, xLLM probably has the | |
| faster attention (FA3), though no one has measured the two side by side. | |
| Neither stack supports AMD GPUs.</em></p> | |
| <h2 id="5-hardware-beyond-nvidia-gpus">5. Hardware beyond NVIDIA | |
| GPUs</h2> | |
| <p><strong>xLLM</strong> is CUDA-only. It hard-codes NCCL, NVML | |
| monitoring and <code>.cuda()</code> calls, loads a compiled CUDA | |
| extension at import, and embeds NVIDIA PTX instructions in its Triton | |
| kernels. There is no ROCm, TPU, Intel or Apple path.</p> | |
| <p><strong>Levanter</strong> runs on TPU v4, v5e, v5p and v6e. TPU tests | |
| run on v5e for pull requests that touch the libraries (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/.github/workflows/unified-unit.yaml#L231-L294" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/.github/workflows/unified-unit.yaml#L231-L294">workflow</a>), | |
| a v6e canary trains daily (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/.github/workflows/marin-canary-ferry.yaml#L3-L5" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/.github/workflows/marin-canary-ferry.yaml#L3-L5">canary</a>), | |
| and Marin trained its 8B, 32B and Snowball 67B-A2B models on TPUs. TPUs | |
| are now secondary for Marin: Will Held wrote on 2026-09-26, <a | |
| href="https://discord.com/channels/1354881461060243556/1553196386843893771/1553196662199947285">"We | |
| are mostly no longer on TPUs!"</a>. The hero's expert-parallel | |
| transport, FA4 and pipeline parallelism have no TPU versions, Marin has | |
| no TPU v7 plan, and on AMD, <a | |
| href="https://github.com/marin-community/marin/issues/9462">"Nobody has | |
| run Levanter or Grug on AMD GPUs."</a></p> | |
| <p><em>Verdict: Levanter, the only one of the two that runs anywhere but | |
| NVIDIA.</em></p> | |
| <h2 id="6-parallelism">6. Parallelism</h2> | |
| <table> | |
| <thead> | |
| <tr> | |
| <th></th> | |
| <th>xLLM</th> | |
| <th>Levanter + Grug</th> | |
| </tr> | |
| </thead> | |
| <tbody> | |
| <tr> | |
| <td>Data parallel / FSDP</td> | |
| <td>FSDP1 or FSDP2 (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/build.py#L73-L150" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/build.py#L73-L150">build.py</a>)</td> | |
| <td>FSDP within a rack or slice; plain data parallel across racks or TPU | |
| slices</td> | |
| </tr> | |
| <tr> | |
| <td>Tensor parallel</td> | |
| <td>Megatron-style; up to 8 ways in practice</td> | |
| <td>Available in classic; unused in the hero by choice (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/6367#issuecomment-4888720086" href="https://github.com/marin-community/marin/issues/6367#issuecomment-4888720086">#6367</a>)</td> | |
| </tr> | |
| <tr> | |
| <td>Context parallel</td> | |
| <td>Gathers all keys and values; no public example or test</td> | |
| <td>Gathers all keys and values; 262K-token tests on GPU (<a | |
| href="https://github.com/marin-community/marin/pull/9119">#9119</a>); | |
| production on TPU</td> | |
| </tr> | |
| <tr> | |
| <td>Expert parallel</td> | |
| <td>Uses the tensor-parallel group, so at most 8 ranks within one | |
| node</td> | |
| <td>Its own mesh axis; 64 ranks across a GB200 rack in production</td> | |
| </tr> | |
| <tr> | |
| <td>Pipeline parallel</td> | |
| <td>Absent</td> | |
| <td>JaxPP benchmarks only (<a | |
| href="https://github.com/marin-community/marin/pull/8739">#8739</a>, <a | |
| href="https://github.com/marin-community/marin/issues/9277">#9277</a>)</td> | |
| </tr> | |
| <tr> | |
| <td>Gradient accumulation</td> | |
| <td>Absent</td> | |
| <td>Classic yes; Grug no</td> | |
| </tr> | |
| <tr> | |
| <td>Largest run</td> | |
| <td>About 600 GPUs for the 7B, inferred from logs; about 2,048 for the | |
| 375B, inferred from a run name</td> | |
| <td>704 GB200s in production since August</td> | |
| </tr> | |
| </tbody> | |
| </table> | |
| <p>In xLLM, each tensor-parallel rank holds a slice of every token's | |
| hidden vector. An all-to-all swaps those slices so that each rank sees | |
| whole vectors for its own experts (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/moe.py#L165-L195" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/moe.py#L165-L195">moe.py</a>). | |
| Tokens never move between data-parallel ranks, so expert parallelism | |
| cannot exceed the tensor-parallel degree, in practice eight GPUs. On | |
| 8-GPU H200 nodes that is a common layout, and the 375B run name says | |
| <code>ep8</code>. But FSDP must then gather each rank's full set of | |
| local experts at every layer, and the design cannot use a 72-GPU NVLink | |
| domain. Marin measured this choice on one GB200 rack: at matched token | |
| drops, expert parallelism reached 22.90% MFU against 19.40% for FSDP at | |
| the same shape (<a | |
| href="https://github.com/marin-community/marin/pull/7981">#7981</a>).</p> | |
| <p>Levanter's pipeline and context parallelism work in benchmarks but | |
| not yet in GPU production; drafts extend both (<a | |
| href="https://github.com/marin-community/marin/pull/9279">#9279</a>, <a | |
| href="https://github.com/marin-community/marin/pull/9460">#9460</a>, <a | |
| href="https://github.com/marin-community/marin/pull/9282">#9282</a>). | |
| The Grug hero has no gradient accumulation, so its batch size is tied to | |
| the number of racks, and its expert parallelism stops at the rack | |
| boundary. On H100s, Marin also uses 8-way expert parallelism within a | |
| node, as xLLM does.</p> | |
| <p><em>Verdict: Levanter for large MoE models. For dense models on 8-GPU | |
| nodes, xLLM's FSDP-first design has proven sufficient.</em></p> | |
| <h2 id="7-mixture-of-experts-support-and-the-375b-model">7. | |
| Mixture-of-experts support and the 375B model</h2> | |
| <p><strong>xLLM supports sparse MoE.</strong> It ships presets for | |
| Mixtral-8x7B, K2 Horizon MoE-36B-A4B, MoVA-36B-A4B and the 375B-A23B (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L877-L898" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L877-L898">presets</a>). | |
| A router picks the top k experts from sigmoid or softmax scores. A | |
| per-expert bias, nudged toward balanced loads at every step, shifts only | |
| which experts are chosen, and an optional <code>dot</code> or | |
| <code>entropy</code> balancing loss can be added (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/router.py#L159-L195" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/router.py#L159-L195">router.py</a>). | |
| Shared experts and leading dense layers are supported, and routing drops | |
| no tokens. Experts run through one of four grouped-GEMM backends, | |
| including a Transformer Engine grouped GEMM spread over several CUDA | |
| streams (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/fused_ops/mgmm/te_mgmm.py#L23-L60" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/fused_ops/mgmm/te_mgmm.py#L23-L60">te_mgmm.py</a>), | |
| and a fused MoE block with a hand-written backward is optional. MoVA | |
| also routes the attention value projections to experts.</p> | |
| <p>The limits matter at frontier scale:</p> | |
| <ul> | |
| <li>Expert parallelism uses the tensor-parallel group, so it spans at | |
| most 8 GPUs in one node (section 6).</li> | |
| <li>Every MoE layer stops to copy expert counts to the host, in both the | |
| eager and fused paths (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/moe.py#L177-L178" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/moe.py#L177-L178">moe.py</a>, | |
| <a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/fused_blocks/moe.py#L90-L91" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/fused_blocks/moe.py#L90-L91">fused | |
| block</a>).</li> | |
| <li>There are no capacity controls and no group-limited routing, and the | |
| router z-loss setting is never read.</li> | |
| <li>Tests cover only the grouped-GEMM and permutation kernels. The | |
| router test checks a copy of the router, and nothing tests a whole MoE | |
| layer or expert-parallel correctness (<a | |
| href="https://github.com/ifm-ai/xllm/tree/889db388d021c92164e7a675f6884dd40ea117c7/tests/moe">tests/moe</a>).</li> | |
| <li>Gekko layers have no MoE variant (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/gekko.py#L345-L348" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/gekko.py#L345-L348">gekko.py</a>).</li> | |
| <li>The only published MoE throughput is MoVA-36B-A4B at a claimed 26.7% | |
| MFU on 128 H200s.</li> | |
| </ul> | |
| <p>Grug, for comparison, offers eight MoE backends (five of them | |
| expert-parallel), capacity accounting, balancing by per-expert quantile | |
| thresholds, and the 64-way expert parallelism across a GB200 rack that | |
| the hero runs (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/grug/_moe/common.py#L87-L111" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/grug/_moe/common.py#L87-L111">common.py</a>, | |
| <a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/grug/grug_moe.py#L188-L344" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/grug/grug_moe.py#L188-L344">grug_moe.py</a>).</p> | |
| <p><strong>Did xLLM train the 375B?</strong> Probably an internal | |
| version of it, but the public record does not prove it.</p> | |
| <p>Evidence for:</p> | |
| <ul> | |
| <li>IFM says so. The <a | |
| href="https://huggingface.co/IFM/K2-Horizon-375B-A23B/blob/82af3bc5fbdbc24f1d15a081ac035396f9b2c3dd/README.md">model | |
| card</a> lists <code>ifm-ai/xllm</code> as its code repository and | |
| advises: <a | |
| href="https://huggingface.co/IFM/K2-Horizon-375B-A23B/blob/82af3bc5fbdbc24f1d15a081ac035396f9b2c3dd/README.md">"To | |
| continue the original pretraining with xLLM's training behavior, use the | |
| native xLLM checkpoint and XLLM runtime."</a></li> | |
| <li>The public preset matches the released <a | |
| href="https://huggingface.co/IFM/K2-Horizon-375B-A23B/blob/82af3bc5fbdbc24f1d15a081ac035396f9b2c3dd/config.json">config</a> | |
| in layer count, leading dense layers, width, attention heads, expert | |
| count and size, sigmoid router with bias and 2.5 scaling, and norm | |
| groups. Only the RoPE base differs, 500K in the preset against 10M in | |
| the release, which fits a change during the move to 512K context.</li> | |
| <li>The run ID for Pretraining Phase 2, | |
| <code>k2moe375B_txt360v2.3_256nodes_seed42_bsz32M_seq8k_jais250k_ep8_dot_te_phase2</code>, | |
| uses xLLM's own option values. <code>dot</code> is a value of | |
| <code>moe_router_load_balancing_type</code> and <code>te</code> of | |
| <code>moe_expert_backend</code> (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L421-L422" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L421-L422">config.py</a>, | |
| <a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L197-L198" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L197-L198">config.py</a>); | |
| Megatron-LM's balancing options include no <code>dot</code> (<a | |
| data-orig-href="https://github.com/NVIDIA/Megatron-LM/blob/73813c854b70afa0ba77ebc76fae8950213ecb20/megatron/training/arguments.py#L3748-L3749" href="https://github.com/NVIDIA/Megatron-LM/blob/73813c854b70afa0ba77ebc76fae8950213ecb20/megatron/training/arguments.py#L3748-L3749">arguments.py</a>). | |
| <code>ep8</code> fits xLLM's tensor-parallel expert layout, which the | |
| preset permits (8 key/value heads, one norm group), and 256 nodes pass | |
| xLLM's startup check.</li> | |
| <li>The 375B's rebuilt <a | |
| href="https://wandb.ai/llm360/K2-Horizon-375B">W&B curves</a> use | |
| xLLM's metric names, <code>optim/avg_loss</code>, | |
| <code>optim/g_norm</code> and <code>optim/lr</code> (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/train.py#L480-L499" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/train.py#L480-L499">train.py</a>). | |
| The Phase 1 curves came from a folder named | |
| <code>phase1_xllm_eval_tag_inventory/tb</code>, the kind of | |
| <code>tb</code> directory xLLM writes (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/metrics.py#L40" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/metrics.py#L40">metrics.py</a>).</li> | |
| <li>The released model code says it returns <a | |
| href="https://huggingface.co/IFM/K2-Horizon-375B-A23B/blob/82af3bc5fbdbc24f1d15a081ac035396f9b2c3dd/modeling_k2_horizon.py">"native-XLLM-compatible | |
| routing weights"</a>, and the config follows the schema of xbridges' | |
| xLLM-to-Hugging Face converter.</li> | |
| </ul> | |
| <p>Gaps:</p> | |
| <ul> | |
| <li>The 375B's W&B project holds no training telemetry. Its eight | |
| runs were rebuilt after the fact from an evaluation spreadsheet and a | |
| TensorBoard export, and record only loss, gradient norm, learning rate | |
| and evaluations. Unlike the 0.9B, 3.7B and 7B logs, they carry no | |
| throughput, token counter or configuration, so the run's speed and | |
| hardware are unknown; <code>256nodes</code> in the run name implies | |
| about 2,048 GPUs.</li> | |
| <li>No native xLLM checkpoint, 375B launch script or technical report is | |
| public.</li> | |
| <li>The released config does not exactly match the public converter's | |
| output: it lacks the converter's <code>xllm_model_parallel_size</code> | |
| field and records bfloat16 where the converter writes float32 (<a | |
| data-orig-href="https://github.com/ifm-ai/xbridges/blob/227f3fd26e30b024ce09bc7c556c64616c13276e/xbridges/huggingface/xllm_to_hf_main.py#L185-L187" href="https://github.com/ifm-ai/xbridges/blob/227f3fd26e30b024ce09bc7c556c64616c13276e/xbridges/huggingface/xllm_to_hf_main.py#L185-L187">xllm_to_hf_main.py</a>).</li> | |
| <li>IFM had Megatron at hand: it pretrained its previous flagship, K2-V2 | |
| 70B, with Megatron-Core (<a | |
| href="https://arxiv.org/abs/2512.06201">K2-V2</a>), and it runs | |
| reinforcement learning on Megatron-LM.</li> | |
| </ul> | |
| <p><em>Verdict: xLLM supports MoE, and an internal version of it | |
| probably pretrained the 375B-A23B on about 2,048 GPUs for 15T tokens. | |
| Its MoE design suits 8-GPU nodes but cannot use a larger NVLink domain, | |
| and IFM has not published how efficiently it trained at that size. | |
| Levanter's hero remains the only published MoE throughput at frontier | |
| scale.</em></p> | |
| <h2 id="8-operations">8. Operations</h2> | |
| <ul> | |
| <li><strong>Scheduler and storage.</strong> xLLM launches under torchrun | |
| or Slurm, and its example scripts, requeue handler and asynchronous | |
| evaluation assume Slurm (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/distributed/slurm.py#L16-L48" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/distributed/slurm.py#L16-L48">slurm.py</a>). | |
| It reads and writes a shared POSIX filesystem and has no object-store | |
| support. Marin runs Levanter through its own Iris scheduler on | |
| Kubernetes and TPU pools, and writes checkpoints with TensorStore to | |
| object storage (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/tensorstore_serialization.py#L942-L1040" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/tensorstore_serialization.py#L942-L1040">tensorstore_serialization.py</a>). | |
| Levanter also detects Slurm jobs (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/distributed.py#L27-L31" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/distributed.py#L27-L31">distributed.py</a>).</li> | |
| <li><strong>Checkpoint and resume.</strong> xLLM saves optimizer state | |
| per rank by default, so a resume must keep the same layout: a new | |
| tensor-parallel degree raises <code>NotImplementedError</code> (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/reloading.py#L150-L162" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/reloading.py#L150-L162">reloading.py</a>), | |
| and a new data-parallel size fails in the data iterator (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/data/dataset_streamer/data_iterator/data_iterator.py#L189-L191" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/data/dataset_streamer/data_iterator/data_iterator.py#L189-L191">data_iterator.py</a>). | |
| Turning on its asynchronous checkpointer fails at startup, because the | |
| class leaves abstract methods unimplemented (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/checkpointing.py#L105-L132" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/checkpointing.py#L105-L132">checkpointing.py</a>). | |
| Levanter writes asynchronously and restores any checkpoint onto any | |
| mesh.</li> | |
| <li><strong>Fault tolerance.</strong> In the public release, Slurm's | |
| warning signal makes xLLM requeue the job without saving, and the only | |
| hang detection is PyTorch's NCCL watchdog. IFM's production runs had | |
| more. Their logs record GPU hardware-error checks and NCCL benchmarks | |
| that the public code cannot produce, and IFM's K2-V2 paper describes an | |
| in-house tool that detects hardware faults and recovers automatically | |
| (<a href="https://arxiv.org/abs/2512.06201">K2-V2</a>). The public code | |
| also differs in small ways from what ran: its three 38-node example | |
| scripts enable a startup bandwidth check that requires a multiple of | |
| four nodes, so as published they stop before training (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/cluster_check/comms_bench.py#L35-L58" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/cluster_check/comms_bench.py#L35-L58">check</a>, | |
| <a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/k2-horizon-7B.sh#L7" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/k2-horizon-7B.sh#L7">script</a>). | |
| Levanter adds a step watchdog (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/callbacks/progress_watchdog.py#L33" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/callbacks/progress_watchdog.py#L33">progress_watchdog.py</a>), | |
| an endpoint that forces a checkpoint, hourly checkpoints on the hero and | |
| automatic retries, so a failure costs the hero at most about an hour of | |
| work. Neither saves a checkpoint when warned of preemption.</li> | |
| <li><strong>Data.</strong> xLLM tokenizes JSONL during training and | |
| packs documents by best fit, so it needs no preprocessing step. Levanter | |
| reads pre-tokenized caches with staged mixtures.</li> | |
| <li><strong>MoE routing.</strong> xLLM drops no tokens, but synchronizes | |
| with the host at every MoE layer. Grug drops assignments above a | |
| capacity factor of 1.15: about 0.01% on the hero at 4K tokens, though in | |
| earlier tests drops climbed from about 7% to 40% when sequences grew | |
| from 4K to 65K tokens (<a | |
| href="https://github.com/marin-community/marin/issues/8435">#8435</a>). | |
| The hero trains with a logit z-loss; xLLM has none.</li> | |
| <li><strong>Determinism.</strong> Levanter claims bitwise determinism on | |
| TPU, even across preemption and resume (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/README.md#L43" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/README.md#L43">README</a>). | |
| Neither stack promises it on GPUs. xLLM's <code>deterministic</code> | |
| flag is off by default and does not reach its Triton MoE kernel, which | |
| adds router gradients with atomic operations (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/permute/permute_kernels.py#L103-L108" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/permute/permute_kernels.py#L103-L108">permute_kernels.py</a>).</li> | |
| </ul> | |
| <p><em>Verdict: Levanter.</em></p> | |
| <h2 id="9-performance">9. Performance</h2> | |
| <table> | |
| <thead> | |
| <tr> | |
| <th>Stack</th> | |
| <th>Model</th> | |
| <th>Hardware</th> | |
| <th>Seq.</th> | |
| <th>MFU</th> | |
| <th>Status</th> | |
| </tr> | |
| </thead> | |
| <tbody> | |
| <tr> | |
| <td>xLLM</td> | |
| <td>K2 Horizon 7B</td> | |
| <td>About 600 GPUs, likely H200 (inferred)</td> | |
| <td>8K</td> | |
| <td>≈48.6% (median 8,735 tokens/s/GPU)</td> | |
| <td>Production, 21.9T tokens (<a | |
| href="https://wandb.ai/llm360/K2-Horizon-7B">logs</a>)</td> | |
| </tr> | |
| <tr> | |
| <td>xLLM</td> | |
| <td>K2 Horizon 7B, four long-context phases</td> | |
| <td>Same</td> | |
| <td>512K</td> | |
| <td>33–51% (medians 632–991 tokens/s/GPU)</td> | |
| <td>Production, 558B tokens (same logs)</td> | |
| </tr> | |
| <tr> | |
| <td>xLLM</td> | |
| <td>Llama3-8B</td> | |
| <td>128 H200, FSDP2</td> | |
| <td>8K</td> | |
| <td>43.3% eager; 44.1% fused; 52.1% fused at 109 GiB/GPU</td> | |
| <td>Claimed (<a | |
| href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/benchmark_llama3_h200_20260925.md">file</a>)</td> | |
| </tr> | |
| <tr> | |
| <td>xLLM</td> | |
| <td>MoVA-36B-A4B</td> | |
| <td>128 H200, FSDP 64 × TP 2</td> | |
| <td>8K</td> | |
| <td>26.7%</td> | |
| <td>Claimed (<a | |
| href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/benchmark_k2-horizon_h200_20260925.md">file</a>)</td> | |
| </tr> | |
| <tr> | |
| <td>Grug</td> | |
| <td>535B-A23B hero</td> | |
| <td>704 GB200, EP64 × DP11</td> | |
| <td>4K</td> | |
| <td>26.68% (median)</td> | |
| <td>Production (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/8317#issuecomment-5821773936" href="https://github.com/marin-community/marin/issues/8317#issuecomment-5821773936">#8317</a>)</td> | |
| </tr> | |
| <tr> | |
| <td>Grug</td> | |
| <td>Snowball 67B-A2B</td> | |
| <td>TPU v4-2048</td> | |
| <td>8K</td> | |
| <td>18.6% (full-attention count)</td> | |
| <td>Production, 10T tokens (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/6044#issuecomment-4812848295" href="https://github.com/marin-community/marin/issues/6044#issuecomment-4812848295">#6044</a>)</td> | |
| </tr> | |
| <tr> | |
| <td>Grug</td> | |
| <td>About 45B MoE</td> | |
| <td>64 H100</td> | |
| <td>—</td> | |
| <td>24.3%</td> | |
| <td>Benchmark (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/6979#issuecomment-4972374164" href="https://github.com/marin-community/marin/issues/6979#issuecomment-4972374164">#6979</a>)</td> | |
| </tr> | |
| <tr> | |
| <td>Grug</td> | |
| <td>Snowball 67B-A2B</td> | |
| <td>64 H100, PP8 × EP8</td> | |
| <td>8K</td> | |
| <td>19.5% (full-attention count)</td> | |
| <td>14 synthetic steps (<a | |
| href="https://github.com/marin-community/marin/pull/8739">#8739</a>)</td> | |
| </tr> | |
| <tr> | |
| <td>Levanter</td> | |
| <td>Llama-3.1-8B</td> | |
| <td>TPU v5p-8</td> | |
| <td>1K</td> | |
| <td>48.6–50.3%</td> | |
| <td>Microbenchmark (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/1864#issuecomment-3508681402" href="https://github.com/marin-community/marin/issues/1864#issuecomment-3508681402">#1864</a>)</td> | |
| </tr> | |
| <tr> | |
| <td>Grug</td> | |
| <td>2.7B dense</td> | |
| <td>8 H100</td> | |
| <td>2K</td> | |
| <td>56.2%</td> | |
| <td>29 steps, on a branch (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/6979#issuecomment-4894291032" href="https://github.com/marin-community/marin/issues/6979#issuecomment-4894291032">#6979</a>)</td> | |
| </tr> | |
| </tbody> | |
| </table> | |
| <p>IFM's production medians sit within 94–101% of its benchmark claims | |
| for the same models, so the benchmark file looks representative. Three | |
| cautions apply.</p> | |
| <ol type="1"> | |
| <li><strong>The stacks count FLOPs differently.</strong> xLLM halves | |
| attention for causal masking, leaves out embedding parameters (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L788-L833" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/models/transformer.py#L788-L833">formula</a>), | |
| and credits attention that document masking skips. Levanter's generic | |
| formula counts attention in full (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/utils/flop_utils.py#L39-L52" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/utils/flop_utils.py#L39-L52">flop_utils.py</a>): | |
| on a Llama3-8B shape it gives the same run 1.13× xLLM's MFU at 8K and | |
| 1.8× at 262K. The hero's counter accounts for sliding windows and sits | |
| within about 2% of a causal count at 4K. The two Snowball rows price | |
| every layer as full attention, though most use a 2K window; counted the | |
| hero's way, they would read about 15–16%.</li> | |
| <li><strong>xLLM's TorchTitan comparison mixes hardware.</strong> xLLM's | |
| file says its TorchTitan row (6,514 tokens/s/GPU) comes from | |
| TorchTitan's documentation, but not that TorchTitan measured it in | |
| December 2024 on 128 H100s capped at 500 W, at half the batch size (<a | |
| data-orig-href="https://github.com/pytorch/torchtitan/blob/d43271c3303d6aa311d74ba27fa116be77b89973/benchmarks/llama3_h100_202412_torchtitan.md#L40-L46" href="https://github.com/pytorch/torchtitan/blob/d43271c3303d6aa311d74ba27fa116be77b89973/benchmarks/llama3_h100_202412_torchtitan.md#L40-L46">TorchTitan</a>); | |
| the file's setup says <a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/benchmark_llama3_h200_20260925.md#L7" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/benchmark_llama3_h200_20260925.md#L7">"All | |
| jobs are running 128 H200 GPUs"</a>.</li> | |
| <li><strong>The 52.1% row needs H200 memory.</strong> It uses 109 GiB | |
| per GPU, and the published script turns every recomputation switch off | |
| (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/llama3-8B.sh#L52-L68" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/examples/llama3-8B.sh#L52-L68">script</a>). | |
| Despite its <code>+ recompute</code> label, the row most likely ran | |
| without recomputation.</li> | |
| </ol> | |
| <p><strong>External baselines.</strong> Marin has calibrated Grug | |
| against Megatron. On 8 B200s, Grug reached 90% of Megatron's throughput | |
| (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/6139#issuecomment-4619246391" href="https://github.com/marin-community/marin/issues/6139#issuecomment-4619246391">#6139</a>). | |
| On one GH200, Megatron was about 28% faster on a small MoE (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/4311#issuecomment-4414429663" href="https://github.com/marin-community/marin/issues/4311#issuecomment-4414429663">#4311</a>). | |
| On a GB200 rack, Megatron-Core reached 31–32% MFU with artificially | |
| balanced routing, 26.6% with Grug's operators swapped in (still | |
| balanced), and 14.7% with its stock learned router (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/7668#issuecomment-5104907558" href="https://github.com/marin-community/marin/issues/7668#issuecomment-5104907558">#7668</a>). | |
| The hero reaches 26.7% across 11 racks with trained routing, though at a | |
| different shape.</p> | |
| <p><em>Verdict: no like-for-like data exists. xLLM's dense numbers are | |
| strong and hold up in production; Marin has no comparable dense | |
| measurement. Levanter's MoE figure is the only published one at frontier | |
| scale; IFM has published none for its 375B. One run would settle the | |
| dense question: Llama3-8B at 8K on the same 64–128 H100s or H200s in | |
| both stacks, scored with one formula.</em></p> | |
| <h2 id="10-evolution-by-agents">10. Evolution by agents</h2> | |
| <p><strong>xLLM</strong> is small and written in PyTorch, which coding | |
| agents know best. But it gives an agent little to steer by: no | |
| AGENTS.md, no CI, and only 6 of its 36 test files run cleanly under | |
| pytest. No model module imports without the compiled extension, | |
| FlashAttention and Flash Linear Attention, so an agent without a CUDA | |
| build cannot run a single model test. The paired eager and fused | |
| implementations, with hand-written backward passes and no parity test, | |
| make each architecture change risky.</p> | |
| <p><strong>Levanter</strong> was reshaped for agents. The Marin monorepo | |
| layers instructions in AGENTS.md files and carries 40 agent skills (for | |
| example <code>change-grug</code>, <code>add-pallas-kernel</code> and | |
| <code>deploy-hero-change</code>); Levanter adds contract tests that | |
| trace every variant on CPU, multi-device tests on simulated CPU devices, | |
| and agent-run lint and review. In September, agents wrote 79% of new | |
| issues, and agent accounts made 28% of training-stack commits; people | |
| also run agents under their own names, so the true share is higher. An | |
| agent decoded GPU hang dumps to trace the QuACK deadlock (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/8870#issuecomment-5687716531" href="https://github.com/marin-community/marin/issues/8870#issuecomment-5687716531">#8870</a>) | |
| and ran the hero's code handoffs (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/8506#issuecomment-5804146010" href="https://github.com/marin-community/marin/issues/8506#issuecomment-5804146010">#8506</a>). | |
| Another runs small-model experiments that anyone can steer by commenting | |
| (<a | |
| href="https://github.com/marin-community/marin/issues/9451">#9451</a>). | |
| David Hall noted that <a | |
| href="https://discord.com/channels/1354881461060243556/1462884917292699669/1464707410533941510">"grug | |
| itself was mostly done by codex"</a>.</p> | |
| <p>Agents still cannot validate GPU kernels, the patched transport or | |
| multi-host behavior without hardware. An agent-run hero handoff filled | |
| the 100 TiB storage quota and cost about two hours and 377 retrained | |
| steps (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/8506#issuecomment-5817400840" href="https://github.com/marin-community/marin/issues/8506#issuecomment-5817400840">#8506</a>). | |
| Reviewers complain that agents <a | |
| href="https://github.com/marin-community/marin/issues/8738">"produce | |
| 'slop' in a way and scope well beyond a normal human coder"</a>.</p> | |
| <p><em>Verdict: Levanter, by a wide margin.</em></p> | |
| <h2 id="11-code-size-and-complexity">11. Code size and complexity</h2> | |
| <p>Counts cover whole libraries, including data, evaluation and export | |
| code, and exclude blank and comment lines.</p> | |
| <table> | |
| <thead> | |
| <tr> | |
| <th></th> | |
| <th>xLLM</th> | |
| <th>Levanter + Haliax + Grug</th> | |
| </tr> | |
| </thead> | |
| <tbody> | |
| <tr> | |
| <td>Python</td> | |
| <td>21.5K (including 2.2K Triton)</td> | |
| <td>57.4K + 9.5K + 17.3K</td> | |
| </tr> | |
| <tr> | |
| <td>Own C++/CUDA</td> | |
| <td>14.8K, plus 26.2K in xattn (11.5K of it modified FlashAttention | |
| code)</td> | |
| <td>1.3K (a DeepEP binding)</td> | |
| </tr> | |
| <tr> | |
| <td>Vendored or pinned kernel code</td> | |
| <td>3.4K lines of Transformer Engine</td> | |
| <td>External packages: FA4, QuACK, CUTLASS DSL, patched XLA plugin</td> | |
| </tr> | |
| <tr> | |
| <td>Mean cyclomatic complexity (Python)</td> | |
| <td>3.6</td> | |
| <td>3.0 / 2.6 / 3.2</td> | |
| </tr> | |
| <tr> | |
| <td>Functions with complexity above 20</td> | |
| <td>2.0%</td> | |
| <td>0.7% / 0.4% / 1.4%</td> | |
| </tr> | |
| <tr> | |
| <td>Largest training-loop function</td> | |
| <td><code>main</code>: complexity 63, 461 lines (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/train.py#L200-L660" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/train.py#L200-L660">train.py</a>)</td> | |
| <td>Hero <code>_run_grug_local</code>: complexity 65, about 390 lines | |
| (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/moe_hero_ep/train.py#L933-L1321" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/moe_hero_ep/train.py#L933-L1321">train.py</a>)</td> | |
| </tr> | |
| </tbody> | |
| </table> | |
| <p>Only Gekko uses xattn. xLLM's complexity sits in its twin | |
| implementations, hand-written backward passes and dead code: about half | |
| its CUDA ops are unreachable, and the published Gecko layer's module is | |
| broken. Levanter's is spread across two model styles, nine attention | |
| backends, eight MoE transports, copied variants and external runtime | |
| pins.</p> | |
| <p><em>Verdict: mixed. xLLM is smaller; Levanter's functions are | |
| simpler, but it has more parts, many outside the repository.</em></p> | |
| <h2 id="12-test-coverage">12. Test coverage</h2> | |
| <p><strong>xLLM</strong> has 62 test functions. About 17 files compare a | |
| kernel's forward and backward passes against a reference, 10 only time | |
| kernels, and three CPU files (23 functions) cover document packing, data | |
| validation and data-loader resume. Nothing tests FSDP, tensor, context | |
| or expert parallelism, checkpoint reloading, MoE layers, the optimizer | |
| or end-to-end loss. The router test checks a local copy of the router | |
| that has already drifted from the production code (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/tests/moe/test_router.py#L1-L7" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/tests/moe/test_router.py#L1-L7">test</a>, | |
| <a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/router.py#L98-L99" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/modules/moe/router.py#L98-L99">production</a>), | |
| and the resume smoke test passes flags that no longer exist. There is no | |
| CI.</p> | |
| <p><strong>Levanter</strong> has about 1,300 test functions and Haliax | |
| 340, and Marin's root suite adds about 230 in files that exercise Grug | |
| or Snowball. They include Hugging Face parity tests, | |
| kernel-versus-reference tests and multi-device tests on simulated CPU | |
| devices. Pull requests run CPU shards and a TPU lane; canaries train a | |
| Grug MoE daily on TPU v6e, 8×H100 and TPU multislice. The gap is GPUs on | |
| pull requests: <a | |
| data-orig-href="https://github.com/marin-community/marin/issues/8704#issuecomment-5443649625" href="https://github.com/marin-community/marin/issues/8704#issuecomment-5443649625">"no | |
| pytest marker anywhere selects GPU work"</a>, the TPU lane ignores | |
| changes under <code>experiments/grug</code>, and the hero's transport is | |
| validated by hand.</p> | |
| <p><em>Verdict: Levanter.</em></p> | |
| <h2 id="13-data-curation-and-mixing">13. Data curation and mixing</h2> | |
| <p>IFM's <a | |
| href="https://github.com/ifm-ai/pretraining-data-toolkit/tree/cd93e116a8217c9e9a45e104b0ea28ac653a998f">pretraining-data-toolkit</a> | |
| and Marin's <a | |
| href="https://github.com/marin-community/marin/tree/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/datakit">Datakit</a> | |
| both sort pretraining documents by topic and quality. They differ in how | |
| much of the pipeline they cover.</p> | |
| <table> | |
| <thead> | |
| <tr> | |
| <th>Stage</th> | |
| <th>IFM toolkit</th> | |
| <th>Marin Datakit</th> | |
| </tr> | |
| </thead> | |
| <tbody> | |
| <tr> | |
| <td>Sources</td> | |
| <td>Documents prepared upstream, such as TxT360's Common Crawl text; | |
| adapters for HPLT, S2ORC, MegaMath, PhilPapers and USPTO</td> | |
| <td>152 open datasets in 292 registry entries, nearly all pinned Hugging | |
| Face revisions (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/src/marin/datakit/sources.py#L141-L635" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/src/marin/datakit/sources.py#L141-L635">sources.py</a>)</td> | |
| </tr> | |
| <tr> | |
| <td>Language ID, rule-based filters, PII removal</td> | |
| <td>Absent; TxT360's public pipeline has language ID and filters for | |
| Common Crawl</td> | |
| <td>Absent as stages; language comes from upstream metadata</td> | |
| </tr> | |
| <tr> | |
| <td>Deduplication</td> | |
| <td>Absent; reads TxT360's duplicate counts, then drops them</td> | |
| <td>Exact by content hash, then MinHash with LSH over character 5-grams | |
| (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/src/marin/processing/classification/deduplication/fuzzy_minhash.py#L145-L172" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/src/marin/processing/classification/deduplication/fuzzy_minhash.py#L145-L172">fuzzy_minhash.py</a>) | |
| and a containment check before removal (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/src/marin/processing/classification/deduplication/cluster_dedup.py#L33-L53" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/src/marin/processing/classification/deduplication/cluster_dedup.py#L33-L53">cluster_dedup.py</a>); | |
| fuzzy removal skips 16 mostly synthetic sources</td> | |
| </tr> | |
| <tr> | |
| <td>Decontamination</td> | |
| <td>Absent</td> | |
| <td>13-gram Bloom filter against the nine Artificial Analysis | |
| Intelligence Index benchmarks and lm-eval tasks (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/src/marin/datakit/decon.py#L102-L130" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/marin/src/marin/datakit/decon.py#L102-L130">decon.py</a>)</td> | |
| </tr> | |
| <tr> | |
| <td>Quality</td> | |
| <td>Three public classifiers (FineWeb-Edu, DCLM, PreSelect); each score | |
| is cut into 20 levels, and the highest sets one of five categories</td> | |
| <td>A small classifier trained on LLM labels; five bands, used for | |
| mixing rather than filtering</td> | |
| </tr> | |
| <tr> | |
| <td>Topic</td> | |
| <td>WebOrganizer's public topic and format labels, kept as columns</td> | |
| <td>Embeddings from Microsoft's Harrier model, clustered into 5,000 | |
| groups and merged into 40 topics</td> | |
| </tr> | |
| <tr> | |
| <td>Output</td> | |
| <td>Shuffled JSONL for each quality category</td> | |
| <td>200 tokenized Levanter caches, one per topic and band</td> | |
| </tr> | |
| <tr> | |
| <td>Mixing</td> | |
| <td>Weights passed to xLLM at launch (<a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L319" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/config.py#L319">config.py</a>); | |
| none published</td> | |
| <td>Staged mixtures in Levanter; weights from proxy-model swarms, | |
| published; search code off <code>main</code></td> | |
| </tr> | |
| <tr> | |
| <td>Orchestration</td> | |
| <td>Scripts run by hand, locally or as Slurm arrays; labeling needs | |
| GPUs</td> | |
| <td>Steps keyed by a hash of their inputs, run on Marin's Zephyr | |
| library, which needs Marin's Iris scheduler beyond one machine; daily | |
| and weekly canaries</td> | |
| </tr> | |
| <tr> | |
| <td>Tests</td> | |
| <td>None; no CI</td> | |
| <td>360 test functions in <code>tests/datakit</code>, plus the engine's | |
| and Levanter's data tests</td> | |
| </tr> | |
| <tr> | |
| <td>Size</td> | |
| <td>About 3,100 lines of Python</td> | |
| <td>About 28,900 lines; 44,600 with the engine</td> | |
| </tr> | |
| <tr> | |
| <td>Data released</td> | |
| <td>About 13.5T tokens in five datasets (my rough estimate)</td> | |
| <td>Corpus not rehosted; recipes download each source from its | |
| creator</td> | |
| </tr> | |
| </tbody> | |
| </table> | |
| <p><strong>IFM's toolkit is the middle of a pipeline.</strong> It labels | |
| documents with five public classifiers, counts tokens with the JAIS-13b | |
| tokenizer, sorts documents into quality categories, shuffles them and | |
| inspects the result. The sorting step reuses the recipe that NVIDIA's | |
| Nemotron-CC dataset used to combine classifier scores, including its | |
| category names and boundaries (<a | |
| data-orig-href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/annotate/classify.py#L176-L234" href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/annotate/classify.py#L176-L234">classify.py</a>). | |
| Its 16 output columns match the 770-million-row | |
| <code>web-high-medium</code> subset of <a | |
| href="https://huggingface.co/datasets/IFM/TxT360-v2">TxT360-v2</a> in | |
| name and order, so this code most likely produced that subset. The | |
| public version was repackaged without being run: <a | |
| data-orig-href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/README.md#L7-L10" href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/README.md#L7-L10">"This | |
| environment could not install the data-processing dependencies or run | |
| CUDA inference"</a>. An agent ran its CPU stages on synthetic data: the | |
| shuffle kept every row and reran identically, and the category shares | |
| came out as expected. Reading the code turns up four problems:</p> | |
| <ul> | |
| <li>WebOrganizer inputs are padded to 8,192 tokens by default, so by my | |
| estimate a 1,000-token page costs about 16 times the compute it needs | |
| (<a | |
| data-orig-href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/annotate/annotate_v2.py#L1060-L1075" href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/annotate/annotate_v2.py#L1060-L1075">annotate_v2.py</a>).</li> | |
| <li>After a model worker fails, the loop moves on to the next file | |
| without clearing queued results, so later files can fail or, rarely, | |
| receive another file's labels (<a | |
| data-orig-href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/annotate/annotate_v2.py#L1185-L1196" href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/annotate/annotate_v2.py#L1185-L1196">annotate_v2.py</a>).</li> | |
| <li>Sorting drops URLs, document IDs and TxT360's duplicate counts, so | |
| released rows cannot be traced to their sources or upsampled by | |
| duplicate count, as IFM did for K2-V2.</li> | |
| <li>The shuffle writes <code>part-*.jsonl</code> files, but xLLM's | |
| loader reads only file names containing <code>chunk</code> and a number | |
| (<a | |
| data-orig-href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/shuffle/shuffle_2ndpass.py#L44" href="https://github.com/ifm-ai/pretraining-data-toolkit/blob/cd93e116a8217c9e9a45e104b0ea28ac653a998f/shuffle/shuffle_2ndpass.py#L44">toolkit</a>, | |
| <a | |
| data-orig-href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/data/dataset_streamer/data_iterator/data_iterator.py#L100-L101" href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/data/dataset_streamer/data_iterator/data_iterator.py#L100-L101">xLLM</a>).</li> | |
| </ul> | |
| <p>For K2 Horizon's newer web sources, such as ClueWeb22 and HPLT, | |
| nothing before labeling is public: no extraction, language ID, | |
| deduplication or decontamination. Its Common Crawl text is older. It | |
| comes from <a | |
| href="https://github.com/LLM360/TxT360/tree/07d98df82292651c2cf56753f9563338ab0c6073">TxT360</a>, | |
| whose extraction, fastText language ID, filtering and global MinHash | |
| deduplication IFM published in 2024 under its former name, LLM360, and | |
| documented in its <a href="https://arxiv.org/abs/2512.06201">K2-V2 | |
| paper</a>. Other IFM repositories cover parts of synthesis and data | |
| checking. <a | |
| href="https://github.com/ifm-ai/search360/tree/4f9a64d47e50cdeec7ecf9c78d2fdfe4d4fcd594">search360</a> | |
| is a retrieval service over <a | |
| data-orig-href="https://github.com/ifm-ai/search360/blob/4f9a64d47e50cdeec7ecf9c78d2fdfe4d4fcd594/README.md#L209-L210" href="https://github.com/ifm-ai/search360/blob/4f9a64d47e50cdeec7ecf9c78d2fdfe4d4fcd594/README.md#L209-L210">1.6 | |
| billion passages</a> that <a | |
| data-orig-href="https://github.com/ifm-ai/search360/blob/4f9a64d47e50cdeec7ecf9c78d2fdfe4d4fcd594/README.md#L16-L18" href="https://github.com/ifm-ai/search360/blob/4f9a64d47e50cdeec7ecf9c78d2fdfe4d4fcd594/README.md#L16-L18">"was | |
| used to generate synthetic training data at billion-token scale"</a> for | |
| K2 Horizon. <a | |
| href="https://github.com/ifm-ai/LCQA/tree/1dee2447ef8c7ea5dddb7d0c2f4122f8bdcb5ede">LCQA</a> | |
| and <a | |
| href="https://github.com/ifm-ai/PRism-synthesis/tree/7fd0264ffe61cbeb309f580bb70fd55111ca996f">PRism-synthesis</a> | |
| generate long-context questions and agentic code data. And the README of | |
| <a | |
| href="https://github.com/ifm-ai/process_entry/tree/a4745b10f5ce71e15d46021301e1ee717dc09bdc">process_entry</a>, | |
| a validator for chat and agent trajectories, says <a | |
| data-orig-href="https://github.com/ifm-ai/process_entry/blob/a4745b10f5ce71e15d46021301e1ee717dc09bdc/README.md#L3-L4" href="https://github.com/ifm-ai/process_entry/blob/a4745b10f5ce71e15d46021301e1ee717dc09bdc/README.md#L3-L4">"Every | |
| conversation used in K2 Horizon mid- and post-training had to pass | |
| it."</a> The generators behind most of the synthetic pretraining data | |
| are not public, and neither is the mixture search.</p> | |
| <p><strong>IFM's view of data.</strong> Mikhail Yurochkin, who leads the | |
| LLM data team at IFM's Silicon Valley lab, set out its approach in a | |
| talk, <a | |
| href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf"><em>Diversity | |
| First</em></a> (undated; the file name says April 9). Three of its | |
| points bear on Marin:</p> | |
| <ul> | |
| <li><em>Mixture search paid little.</em> IFM fit a model that predicts | |
| evaluation scores from mixture weights over sources defined by | |
| WebOrganizer's labels, then sampled and refined candidate mixes. At 1.5B | |
| parameters and 0.6T tokens, the searched mix scored about two points of | |
| MMLU-CoT above a hand-tuned mix, reading the slide's chart. The slide's | |
| title is <a | |
| data-orig-href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=15" href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=15">"It | |
| worked but improvement is Small"</a>, and it notes that the <a | |
| data-orig-href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=15" href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=15">"data | |
| mixing experiment required 375k GPU-hours"</a>. On BBH, the best of four | |
| mixes at 1.5B was the worst at 7B (<a | |
| data-orig-href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=11" href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=11">slide | |
| 11</a>).</li> | |
| <li><em>Base-model benchmarks are the wrong target.</em> Reusing figures | |
| from the <a href="https://arxiv.org/abs/2510.24397">APTBench</a> paper, | |
| the talk shows base-model MMLU correlating weakly (r = 0.38) with | |
| post-trained SWE-bench Verified scores. It proposes scoring mixes by | |
| multiple-choice versions of post-training tasks, by pass@k, or by pass | |
| rates after chat-formatted data is added to pretraining (<a | |
| data-orig-href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=4" href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=4">slides | |
| 4–9</a>).</li> | |
| <li><em>Diversity pays, if synthetic data is grounded in real data.</em> | |
| The recap sets <a | |
| data-orig-href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=16" href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=16">"Data | |
| Mixing: many nuances; computationally expensive; small gains"</a> | |
| against <a | |
| data-orig-href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=16" href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=16">"Data | |
| Diversity: steady and reliable gains"</a>. By IFM's compression measure, | |
| reasoning text generated from seed queries grew nearly as repetitive as | |
| web code; adding retrieval from an index of the pretraining corpus, and | |
| randomness in the prompts, brought it close to high-quality web text (<a | |
| data-orig-href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=24" href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf#page=24">slide | |
| 24</a>).</li> | |
| </ul> | |
| <p>K2 Horizon applied this at scale: about <a | |
| href="https://web.archive.org/web/20260913112353/https://ifm.ai/blog/k2/">"10 | |
| trillion synthetic tokens during pre-training"</a>, roughly half its | |
| pretraining tokens, generated with <a | |
| href="https://web.archive.org/web/20260913112353/https://ifm.ai/blog/k2/">"millions | |
| of combinations of diversity knobs and context seeds, including | |
| retrieval from an internally built search engine over the pre-training | |
| web corpus"</a>. IFM has released five K2 Horizon datasets: <a | |
| href="https://huggingface.co/datasets/IFM/TxT360-v2">TxT360-v2</a> (web | |
| text, some with question-and-answer additions), <a | |
| href="https://huggingface.co/datasets/IFM/Pretrain-Behaviors">Pretrain-Behaviors</a>, | |
| <a | |
| href="https://huggingface.co/datasets/IFM/Math-Reasoning">Math-Reasoning</a>, | |
| <a | |
| href="https://huggingface.co/datasets/IFM/Code-Reasoning">Code-Reasoning</a> | |
| and <a | |
| href="https://huggingface.co/datasets/IFM/SFT-Reasoning">SFT-Reasoning</a>. | |
| By my rough estimate from sampled rows they hold about 13.5T tokens, | |
| most of them synthetic. That is far more pretraining data than Marin has | |
| released, but less than IFM's product page promises (<a | |
| href="https://web.archive.org/web/20260906023511/https://ifm.ai/k2/">"We | |
| provide the full pre-training corpus"</a>):</p> | |
| <ul> | |
| <li>The dataset cards name no generator models or prompts, so users | |
| cannot check license terms inherited from the generators. They say only | |
| that subsets <a | |
| href="https://huggingface.co/datasets/IFM/TxT360-v2">"may have | |
| undergone"</a> deduplication, and never mention decontamination.</li> | |
| <li>No stage's mixture weights are published, although IFM published | |
| exact weights for its previous model (<a | |
| href="https://github.com/LLM360/k2v2_train">k2v2_train</a>).</li> | |
| <li>The pretraining and mid-training datasets named in the model cards | |
| are not public.</li> | |
| <li>Plain web text covers only the Medium-High quality category, the one | |
| subset that carries IFM's quality and topic labels. High-category pages | |
| come with programmatic questions and answers appended, and the | |
| <code>txt360-qa</code> subset, about half of TxT360-v2's tokens, appears | |
| to be K2-V2-era data released again.</li> | |
| <li>In a sampled slice of the <code>web-high-medium</code> subset, about | |
| a tenth of the rows are labeled <code>clueweb</code>. Carnegie Mellon | |
| distributes ClueWeb22 for research only, under signed license agreements | |
| (<a | |
| href="https://web.archive.org/web/20260925071610/https://lemurproject.org/clueweb22/obtain.php">terms</a>), | |
| so users should confirm IFM's right to release those rows under | |
| CC-BY-4.0.</li> | |
| </ul> | |
| <p><strong>Marin's Datakit curates open corpora.</strong> Will Held's <a | |
| href="https://openathena.ai/blog/marin-data-pipeline-overview/">blog | |
| post</a> describes the pipeline. It starts from about 25T tokens of | |
| datasets that others have already collected and cleaned, plus one slice | |
| of Common Crawl that Marin extracted itself. Normalization gives each | |
| document a content hash. Exact and fuzzy deduplication then run across | |
| sources, and a candidate duplicate is removed only if a longer kept | |
| document contains at least 75% of its word 3-grams; the blog reports | |
| 12.5% of documents and 8.4% of tokens removed. Decontamination checked | |
| 18.7 billion documents and marked 259,960 (<a | |
| href="https://github.com/marin-community/marin/pull/8327">#8327</a>). | |
| The hero trains from a 23.1T-token store of the 200 topic-and-band | |
| buckets.</p> | |
| <p>Datakit has gaps. It has no language identification, rule-based | |
| filters or corpus-wide PII removal. Its weekly Nemotron-CC canary has | |
| failed every week since August 17; since September 7 the cause has been | |
| a verification step that cannot get 1,000 workers ready in time (<a | |
| href="https://github.com/marin-community/marin/issues/8948">#8948</a>, | |
| <a | |
| href="https://github.com/marin-community/marin/issues/9494">#9494</a>). | |
| And the code behind the hero's distinctive choices is not on | |
| <code>main</code>. The topic-clustering pull request was closed unmerged | |
| (<a | |
| href="https://github.com/marin-community/marin/pull/8208">#8208</a>), | |
| the quality-scoring one is still open (<a | |
| href="https://github.com/marin-community/marin/pull/8303">#8303</a>), | |
| and both mixture-search ones were closed (<a | |
| href="https://github.com/marin-community/marin/pull/7541">#7541</a>, <a | |
| href="https://github.com/marin-community/marin/pull/8633">#8633</a>). | |
| From <code>main</code>, Marin can rerun the same deduplication and | |
| decontamination rules, but not the steps that assigned the hero's topics | |
| and quality bands or chose its weights. Those survive only as published | |
| artifacts: the <a | |
| href="https://huggingface.co/marin-community/marin-data-mix-tools">topic | |
| centroids and quality model</a> and the <a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/moe_hero_ep/harrier_mix_2026_08_18.py#L36-L49" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/grug/moe_hero_ep/harrier_mix_2026_08_18.py#L36-L49">mixture | |
| spec</a>. Marin does not rehost its corpus, because <a | |
| href="https://discord.com/channels/1354881461060243556/1462895580064911522/1543322065484910652">"several | |
| datasets we use forbid redistribution"</a>; its recipes download each | |
| source from its creator. It does publish data it generates, such as its | |
| <a | |
| href="https://huggingface.co/datasets/marin-community/grug-moe-mix-swarm">proxy-run | |
| results</a> and <a | |
| href="https://huggingface.co/datasets/open-athena/Snowball-67B-A2B-RLVR1-Repro-Data">RL | |
| reproduction data</a>.</p> | |
| <p><strong>How Marin mixes.</strong> Marin searches with swarms of small | |
| proxy MoE models. Each proxy trains on a scaled-down pool, repeating it | |
| as often as the hero would repeat the full one; a regression fit to the | |
| results proposes weights over the 200 buckets, and a ladder of larger | |
| models checks them. The objective is BPB on held-out text and on the | |
| answers of base-model benchmarks. Marin's results echo IFM's:</p> | |
| <ul> | |
| <li>In June, at 3e17–3e19 FLOPs, a curated mix tied proportional | |
| sampling (1.14× compute-equivalent, 90% interval 0.93–1.41), while the | |
| curated mix on Datakit's new data beat the old Nemotron-based mix 1.78× | |
| (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/6757#issuecomment-4847875950" href="https://github.com/marin-community/marin/issues/6757#issuecomment-4847875950">tie</a>, | |
| <a | |
| data-orig-href="https://github.com/marin-community/marin/issues/6757#issuecomment-4847173767" href="https://github.com/marin-community/marin/issues/6757#issuecomment-4847173767">new | |
| data</a>). Will Held concluded that <a | |
| href="https://discord.com/channels/1354881461060243556/1520123383138750595/1520140544926421136">"new | |
| data is right now a far bigger accelerant"</a>.</li> | |
| <li>The hero's August launch mix <a | |
| href="https://discord.com/channels/1354881461060243556/1462895580064911522/1539387191216709673">"looked | |
| to underperform our old mix"</a> on the pre-launch ladder.</li> | |
| <li>A September re-mix beat the August mix 1.20× at the largest ladder | |
| rung, in a single-seed comparison, and has fed the hero since step | |
| 108,000 on 2026-09-15 (<a | |
| href="https://github.com/marin-community/marin/issues/9126">#9126</a>, | |
| <a | |
| data-orig-href="https://github.com/marin-community/marin/issues/8506#issuecomment-5684891157" href="https://github.com/marin-community/marin/issues/8506#issuecomment-5684891157">#8506</a>). | |
| I found no run at scale that compares the final mix with proportional | |
| sampling.</li> | |
| </ul> | |
| <p>The search is not cheap. By my estimate from Marin's records, the | |
| June swarm of 840 proxy runs cost about 370,000 TPU-v4 chip-hours (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/7067#issuecomment-4990793633" href="https://github.com/marin-community/marin/issues/7067#issuecomment-4990793633">#7067</a>), | |
| comparable in scale to IFM's 375,000 GPU-hours, and the September swarm | |
| 27,000–30,000 H100-hours. The blog's own advice is modest: <a | |
| href="https://openathena.ai/blog/marin-data-pipeline-overview/">"When in | |
| doubt, using proportional sampling or UniMax will probably be better | |
| than a poorly executed learned mixture."</a></p> | |
| <p>The labs differ in their target and their synthetic share. Marin | |
| still optimizes BPB, and its tests of post-training proxies were mixed: | |
| base-model loss on math traces predicted post-RL accuracy across ten | |
| models (r = −0.89) but not the gain from RL (R² = 0.33) (<a | |
| href="https://github.com/marin-community/marin/issues/6096">#6096</a>). | |
| About half of K2 Horizon's pretraining tokens were synthetic, while | |
| Marin's synthetic data comes mostly from NVIDIA's Nemotron releases. | |
| Marin's own research found that rephrasing web text gave 1.48× data | |
| efficiency, and 1.80× when rephrasings were stitched into longer | |
| documents, in a data-constrained setting (<a | |
| href="https://github.com/marin-community/marin/issues/3905">#3905</a>). | |
| Datakit has no stage that generates such data.</p> | |
| <p><em>Verdict: Marin for code, IFM for released data. Datakit covers | |
| deduplication, decontamination and tokenization with tests, and Marin | |
| publishes its mixture weights, proxy-run data and ladder results, though | |
| not the code that chose the hero's topics, bands and weights. IFM has | |
| released trillions of tokens, most of them synthetic, but no mixture | |
| weights and little of the code behind its newest data. Both labs found | |
| that mixture search gains less than new or more diverse data.</em></p> | |
| <h2 id="14-post-training">14. Post-training</h2> | |
| <p>Both labs run RL in PyTorch with Megatron as the learner; all ten | |
| open RL frameworks that Marin surveyed in June were PyTorch (<a | |
| href="https://github.com/marin-community/marin/issues/6162">#6162</a>). | |
| IFM runs RL with <a | |
| href="https://github.com/ifm-ai/RL360/tree/5b548b6948e6e78af99307a38ee96057ea571faa">RL360</a> | |
| and supervised fine-tuning (SFT) with xLLM, also PyTorch. Marin runs RL | |
| with <a | |
| href="https://github.com/marin-community/MarinSkyRL/tree/0cdccc9229f47c8eadff90aa2033a5b4604266aa">MarinSkyRL</a> | |
| but fine-tunes in JAX.</p> | |
| <table> | |
| <thead> | |
| <tr> | |
| <th></th> | |
| <th>IFM</th> | |
| <th>Marin</th> | |
| </tr> | |
| </thead> | |
| <tbody> | |
| <tr> | |
| <td>Recipe</td> | |
| <td>Mid-training shifts toward reasoning and agentic data; then RL | |
| experts, a weight merge, and 250–330B tokens of SFT at 512K (3.7B, 7B, | |
| 375B). The 0.9B, mid-trained only to 128K, has no SFT: on-policy | |
| distillation from its experts follows the merge. The 32B and MoVA get | |
| SFT only (269B tokens)</td> | |
| <td>SFT (80% chat data, 20% pretraining replay, 262K context), then RL | |
| with verifiable rewards across mixed domains; experts and distillation | |
| planned</td> | |
| </tr> | |
| <tr> | |
| <td>SFT code</td> | |
| <td>xLLM (PyTorch): chat data in the pretraining loader, loss on | |
| assistant tokens</td> | |
| <td>JAX: a TPU-only copy of the Grug trainer on an unmerged branch, with | |
| loss on all tokens, for the September Datakit cuts; Levanter with | |
| assistant-only loss, from a launcher off <code>main</code>, for a later | |
| cut</td> | |
| </tr> | |
| <tr> | |
| <td>RL framework</td> | |
| <td>Snapshot of IFM's public forks of Miles, Megatron-LM, SGLang, the | |
| SMG router and Harbor</td> | |
| <td>Hard fork of SkyRL; Megatron-Core only since 2026-09-25; Marin's | |
| vLLM fork; a Harbor fork</td> | |
| </tr> | |
| <tr> | |
| <td>Algorithms in the public recipe or configs</td> | |
| <td>GRPO without standard-deviation scaling, DAPO-style clipping and | |
| filtering, TIS, one-step-off asynchronous rollouts (<a | |
| data-orig-href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/recipes/coding-overfit32.sbatch#L111-L124" href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/recipes/coding-overfit32.sbatch#L111-L124">recipe</a>)</td> | |
| <td>GRPO and RLOO-N with PPO-style clipping, TIS, router replay, | |
| bounded-staleness asynchronous training</td> | |
| </tr> | |
| <tr> | |
| <td>Rewards</td> | |
| <td>Harbor task tests; a 9,900-line verifier library, half of it bundled | |
| benchmark code, that no launcher uses</td> | |
| <td>Harbor tasks; skyrl-gym verifiers, including NVIDIA's Nemotron-Ultra | |
| graders and generative reward model; LLM judges</td> | |
| </tr> | |
| <tr> | |
| <td>Combining experts</td> | |
| <td>Weight merges (ISO and RAM for the 3.7B and 7B, task arithmetic for | |
| the 0.9B); code private, but experts, merged checkpoints and settings | |
| released</td> | |
| <td>No weight merging; multi-teacher on-policy distillation planned and | |
| smoke-tested</td> | |
| </tr> | |
| <tr> | |
| <td>Model support</td> | |
| <td>K2 Horizon 7B recipe; 375B and MoVA code paths without recipes; no | |
| export to Hugging Face format</td> | |
| <td>Snowball 67B-A2B through a Megatron port of Grug; hero support in | |
| open pull requests</td> | |
| </tr> | |
| <tr> | |
| <td>Operations</td> | |
| <td>Slurm and Docker; one recipe; no resume</td> | |
| <td>Marin's Iris launcher; checkpoints streamed to S3; x86 hosts only on | |
| <code>main</code>, so no GB200</td> | |
| </tr> | |
| <tr> | |
| <td>Tests and CI</td> | |
| <td>None for IFM's own code in the snapshot; the Miles fork keeps about | |
| 140 test files, but its test workflow is off</td> | |
| <td>About 2,300 test functions; CPU tests on every pull request; three | |
| nightly H100 lanes</td> | |
| </tr> | |
| <tr> | |
| <td>Delivered</td> | |
| <td>Six post-trained K2 Horizon models, 0.9B to 375B; RL experts for | |
| four of them, including the 375B-A23B MoE</td> | |
| <td>RL checkpoints of the 67B-A2B, released as research artifacts; first | |
| official release pushed back a week from October 6</td> | |
| </tr> | |
| </tbody> | |
| </table> | |
| <p><strong>RL360 is a release snapshot.</strong> Its README sums up the | |
| design: <a | |
| data-orig-href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/README.md#L41" href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/README.md#L41">"Miles | |
| manages rollouts and training, Megatron-LM updates the policy, SMG | |
| routes requests to SGLang, and Harbor runs tasks and verifiers in | |
| sandboxes"</a>. Outside its vendored components it holds 14,200 lines, | |
| 1.6% of the tree, and about 5,000 of those are bundled benchmark | |
| verifiers such as IFEval and LiveBench. The rest is vendored from IFM's | |
| public forks, without their Python tests. IFM's work inside those forks | |
| is not in the 14,200: its Miles fork alone is 50 commits and about 9,600 | |
| added lines ahead of upstream (<a | |
| href="https://github.com/LLM360/miles/compare/85fdb7e782388b9584150d72405d09cbd8ad15f4...de89a5ee12d026b52752268813b5c45ac3d9a132">compare</a>). | |
| That work wires Harbor sandboxes into multi-turn rollouts, keeps sandbox | |
| failures out of training, and adds K2 Horizon and MoVA support. The | |
| Megatron-LM copy dates from February 2026 (megatron-core 0.16.0rc0).</p> | |
| <p>The one recipe trains the 7B to overfit 32 coding tasks on 33 | |
| eight-GPU nodes: 64 GPUs train, 192 generate and one node runs services | |
| (<a | |
| data-orig-href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/README.md#L47" href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/README.md#L47">README</a>). | |
| The README warns that <a | |
| data-orig-href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/README.md#L165" href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/README.md#L165">"the | |
| Docker build, checkpoint conversion, and full GPU run still need to be | |
| tested together on the target cluster"</a>. Its <a | |
| href="https://wandb.ai/mbzuai-llm/rl360-public/runs/wm7usgst">public | |
| run</a>, on H200s, hit its 72-hour limit after 99 of 100 planned steps, | |
| about 18,900 GPU-hours. Training reward rose from 0.37 to 0.90, while | |
| the training GPUs sat idle 97% of each 40-minute step as agents worked | |
| in sandboxes. The demo does not show IFM's production pace: the 7B's | |
| shipped code experts logged median idle shares of 12–13% (<a | |
| href="https://wandb.ai/llm360/K2-Horizon-7B/runs/m7tqeo5y">run</a>), | |
| though its math and tool-use experts idled 63% and 88%.</p> | |
| <p>RL360 is also not the exact code behind K2 Horizon's shipped experts. | |
| The 3.7B's and 7B's math, code and STEM experts ran from an internal | |
| RL360 checkout but logged metrics this snapshot cannot produce. Two of | |
| the 7B's four experts <a | |
| href="https://wandb.ai/llm360/K2-Horizon-7B/runs/4blppt0e">"were trained | |
| by another team"</a>; the 0.9B's RL and distillation ran on slime (<a | |
| href="https://wandb.ai/llm360/K2-Horizon-0.9B/runs/2y4evxdr">run</a>); | |
| and the 375B's experts came from a harness named <code>fmp</code> (<a | |
| href="https://wandb.ai/llm360/K2-Horizon-375B/runs/bot25bfe">run</a>). | |
| The 7B's card describes an <a | |
| href="https://huggingface.co/IFM/K2-Horizon-7B">"ISO merge on | |
| self-attention and shared experts, RAM on the remaining weights"</a>, | |
| but no merge code is public, and SFT ran on xLLM (<a | |
| href="https://wandb.ai/llm360/K2-Horizon-7B/runs/wms33y11">run</a>). So | |
| IFM's promise, <a | |
| href="https://web.archive.org/web/20260906023511/https://ifm.ai/k2/">"We | |
| make both our pre-training and post-training code available so you can | |
| rerun or modify the training pipeline"</a>, holds only in part. Reading | |
| the code also turns up three defects:</p> | |
| <ul> | |
| <li>The documented task file lacks a field that the agent code requires, | |
| so rollouts set up as documented would abort (<a | |
| data-orig-href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/src/agent360/harbor/miles/swe_agent_function.py#L110-L133" href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/src/agent360/harbor/miles/swe_agent_function.py#L110-L133">swe_agent_function.py</a>).</li> | |
| <li>The snapshot cannot export a trained K2 Horizon policy to Hugging | |
| Face format: the conversion tool was left out, and | |
| <code>--save-hf</code> swallows its errors (<a | |
| data-orig-href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/components/miles/miles/backends/megatron_utils/model.py#L751-L790" href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/components/miles/miles/backends/megatron_utils/model.py#L751-L790">model.py</a>).</li> | |
| <li>The unused code verifier runs model-written code without a sandbox | |
| by default (<a | |
| data-orig-href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/src/miles360/reward/coder1/__init__.py#L15-L25" href="https://github.com/ifm-ai/RL360/blob/5b548b6948e6e78af99307a38ee96057ea571faa/src/miles360/reward/coder1/__init__.py#L15-L25">coder1</a>).</li> | |
| </ul> | |
| <p><strong>MarinSkyRL is a hard fork that Marin keeps | |
| rewriting.</strong> Benjamin Feuer brought it from the OpenThoughts | |
| project in June, and its policy reads: <a | |
| data-orig-href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/AGENTS.md#L10-L11" href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/AGENTS.md#L10-L11">"This | |
| is a <strong>hard snapshot</strong>: we own this tree and no upstream | |
| sync or merge-back is planned"</a>. About 74% of its active Python lines | |
| were last written in the Marin era, and 628 pull requests have merged | |
| since mid-July, 81% of them Feuer's, so, like the hero's GB200 path, it | |
| leans heavily on one engineer. It also changes fast. On 2026-09-25, <a | |
| href="https://github.com/marin-community/MarinSkyRL/pull/776">#776</a> | |
| deleted the FSDP, DeepSpeed and LoRA code, 36,000 lines, about half of | |
| them tests. Megatron-Core is now the only trainer (<a | |
| data-orig-href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/utils/utils.py#L482-L486" href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/utils/utils.py#L482-L486">utils.py</a>), | |
| and its runtime installs only on x86 hosts, so GB200 support waits on an | |
| open pull request (<a | |
| href="https://github.com/marin-community/MarinSkyRL/pull/798">#798</a>). | |
| On 2026-09-28, <a | |
| href="https://github.com/marin-community/MarinSkyRL/pull/774">#774</a> | |
| replaced the training loop. Three experiments on Marin's | |
| <code>main</code> still request the deleted FSDP2 path (<a | |
| href="https://github.com/marin-community/MarinSkyRL/issues/809">MarinSkyRL | |
| #809</a>). RL trains with AdamW: the Megatron path is set up and tested | |
| only for Adam and SGD (<a | |
| data-orig-href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/distributed/megatron/optimizer.py#L26-L35" href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/distributed/megatron/optimizer.py#L26-L35">optimizer.py</a>), | |
| and the Grug MuonH optimizer went with the FSDP code; MuonH ports are in | |
| draft pull requests (<a | |
| href="https://github.com/marin-community/MarinSkyRL/pull/820">#820</a>). | |
| Both forks lag their upstreams. MarinSkyRL is 815 commits behind SkyRL, | |
| and in August Marin judged a rebase intractable (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/8048#issuecomment-5243952047" href="https://github.com/marin-community/marin/issues/8048#issuecomment-5243952047">"The | |
| rebase is still intractable"</a>); IFM's Miles fork is 1,515 commits | |
| behind Miles.</p> | |
| <p>Rollouts run on Marin's vLLM fork. A weight sync that sends each | |
| expert straight to the engine that serves it installs a Snowball update | |
| in 0.51 s, against 19.9 s for a full broadcast (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/8955#issuecomment-5634012950" href="https://github.com/marin-community/marin/issues/8955#issuecomment-5634012950">#8955</a>), | |
| and EAGLE-3 draft models trained on Snowball's own rollouts made agentic | |
| RL steps 1.42× faster (<a | |
| href="https://github.com/marin-community/marin/issues/9114">#9114</a>). | |
| The SGLang code is present but disabled. The registries hold the common | |
| GRPO-family estimators and nine policy losses (<a | |
| data-orig-href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/utils/algorithm_registry.py#L21-L27" href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/utils/algorithm_registry.py#L21-L27">estimators</a>, | |
| <a | |
| data-orig-href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/utils/algorithm_registry.py#L86-L96" href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/utils/algorithm_registry.py#L86-L96">losses</a>), | |
| alongside TIS, router replay, bounded-staleness asynchronous training | |
| (<a | |
| data-orig-href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/config/ppo_base_config.yaml#L260-L275" href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/skyrl_train/config/ppo_base_config.yaml#L260-L275">config</a>) | |
| and on-policy distillation. There is no critic, and so no GAE-based PPO; | |
| no reward-model or DPO training; and no standalone SFT trainer. About | |
| 2,250 CPU test functions run on every pull request, and the nightly H100 | |
| workflow, now three lanes, passed 62 of its 84 runs since July 15. But | |
| no automated run covers asynchronous training, agentic RL on GPUs or | |
| multi-node Megatron, and the nightly gate checks only that runs finish | |
| with finite metrics.</p> | |
| <p><strong>Crossing from JAX to PyTorch costs Marin.</strong> Snowball | |
| was pretrained in JAX, so RL needs a PyTorch copy of Grug, checkpoint | |
| conversion and parity tests. The first PyTorch port, written as a | |
| correctness baseline, ran 27.3× slower than Levanter per update (<a | |
| href="https://github.com/marin-community/marin/pull/7985">#7985</a>); | |
| grouped expert kernels cut that to 1.38× on the same replay (<a | |
| href="https://github.com/marin-community/MarinSkyRL/issues/259">MarinSkyRL | |
| #259</a>). FSDP2 then failed a gate that required the trainer's token | |
| probabilities to match vLLM's exactly, and Marin kept Megatron (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/8955#issuecomment-5652989341" href="https://github.com/marin-community/marin/issues/8955#issuecomment-5652989341">#8955</a>). | |
| Parity rests on a chain of tiny-model checks: a fixture generated by | |
| Levanter, with hidden size 8 and four experts, checks the PyTorch model | |
| at FP32 tolerances (<a | |
| data-orig-href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/tests/grug_training_parity.py#L15-L19" href="https://github.com/marin-community/MarinSkyRL/blob/0cdccc9229f47c8eadff90aa2033a5b4604266aa/skyrl-train/tests/grug_training_parity.py#L15-L19">test</a>), | |
| and the PyTorch model then checks the Megatron port. No test compares | |
| Levanter with Megatron at Snowball's full width. Marin chose SkyRL in | |
| June partly because Miles and slime were <a | |
| data-orig-href="https://github.com/marin-community/marin/issues/6162#issuecomment-4622595016" href="https://github.com/marin-community/marin/issues/6162#issuecomment-4622595016">"SGLang-only"</a> | |
| and Marin wanted to keep vLLM. It deleted its in-process JAX RL engine | |
| in August (<a | |
| href="https://github.com/marin-community/marin/pull/8063">#8063</a>). | |
| JAX learners now exist only as drafts (<a | |
| href="https://github.com/marin-community/marin/pull/9050">#9050</a>, <a | |
| href="https://github.com/marin-community/marin/pull/9144">#9144</a>), | |
| though David Hall measured a JAX Grug learner at only about 4.6% slower | |
| than Megatron's pipeline-parallel one on H100s (<a | |
| href="https://github.com/marin-community/marin/issues/8970">#8970</a>). | |
| IFM crosses a smaller gap: xLLM and Megatron-LM are both PyTorch, and | |
| IFM's forks add grouped RMSNorm, xLLM's partial-RoPE layout and MoVA to | |
| Megatron, and an xLLM checkpoint bridge to Miles.</p> | |
| <p><strong>Results so far.</strong> Marin's release pick, | |
| <code>Snowball-67B-A2B-10T-Mixed-RLVR-Sync-Step92</code>, came from | |
| mixed-domain RL on 64 H100s per arm, after the September 11 SFT, and was | |
| chosen on the same 26-benchmark panel that reports its scores (<a | |
| href="https://github.com/marin-community/marin/issues/9359">#9359</a>, | |
| <a | |
| href="https://github.com/marin-community/marin/issues/9412">#9412</a>). | |
| It scores 83.8% on MATH-500, 43.9% on AIME24, 43.1% on GPQA-Diamond and | |
| 43.7% on MMLU-Pro, but 3.0% on a 100-task SWE-bench Verified sample and | |
| 1.2% on Terminal-Bench 2. The agentic scores may understate the model: | |
| the later September 21 SFT cut, in its default thinking mode, writes | |
| tool calls that the Terminus-2 agent harness <a | |
| data-orig-href="https://github.com/marin-community/marin/issues/9225#issuecomment-5786970615" href="https://github.com/marin-community/marin/issues/9225#issuecomment-5786970615">"cannot | |
| parse"</a>, and with thinking off it solved 21 of the same 100 SWE-bench | |
| tasks. The second RL stage was <a | |
| href="https://github.com/marin-community/marin/issues/9359">"flat to | |
| regressing"</a>, and pruning deleted the checkpoints at its pass@16 | |
| peaks. On an older lineage, SFT followed by 30 GRPO steps lifted the | |
| SWE-bench sample from 14.5% to 30.7% (<a | |
| data-orig-href="https://github.com/marin-community/marin/issues/9225#issuecomment-5777718831" href="https://github.com/marin-community/marin/issues/9225#issuecomment-5777718831">#9225</a>). | |
| Like IFM's, Marin's results came from code that has since changed: | |
| Step92 predates the new training loop, and the SWE-bench run used the | |
| FSDP2 trainer that #776 deleted.</p> | |
| <p><strong>Fine-tuning.</strong> Marin's SFT code is scattered. The | |
| September 20 and 21 Datakit SFT cuts came from a copy of the Grug | |
| trainer on an unmerged branch, which Will Held called <a | |
| href="https://discord.com/channels/1354881461060243556/1551721992040882276/1552128612260388924">"TPU | |
| specific still"</a> and which computes loss on every token, including | |
| user and tool turns (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/8a0ec6cf352792bf6da963c2a5977f312afdf571/experiments/grug_sft/special_token_lr.py#L94-L100" href="https://github.com/marin-community/marin/blob/8a0ec6cf352792bf6da963c2a5977f312afdf571/experiments/grug_sft/special_token_lr.py#L94-L100">special_token_lr.py</a>). | |
| Asked why, Held answered <a | |
| href="https://discord.com/channels/1354881461060243556/1552024155959066787/1552093562928103527">"No | |
| well founded reason other than that masking them doesn't accelerate | |
| training much"</a>, adding that past studies had found training on user | |
| turns neutral to positive. A September 23 SFT used assistant-only loss | |
| through Levanter, from a launcher not on <code>main</code> (<a | |
| href="https://huggingface.co/open-athena/Grug-67B-A2B-GLM53-RLVR-SFT-2026.09.23">card</a>). | |
| On <code>main</code>, a Grug SFT backend supports masked, packed | |
| training (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/june_tpu_67b_a2b/moe/train.py#L443-L444" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/experiments/june_tpu_67b_a2b/moe/train.py#L443-L444">train.py</a>), | |
| and a newer Levanter path for Snowball masks non-assistant tokens but | |
| cannot pack sequences (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/models/snowball.py#L744-L745" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/models/snowball.py#L744-L745">snowball.py</a>); | |
| its checked-in recipes run 4 and 10 steps. IFM's SFT ran on xLLM, whose | |
| public code trains only on assistant tokens. No experiment on | |
| <code>main</code> uses Levanter's DPO or LoRA code, and its LoRA wraps | |
| only Haliax <code>Linear</code> layers, which Grug models lack (<a | |
| data-orig-href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/adaptor/lora.py#L158-L162" href="https://github.com/marin-community/marin/blob/4aa26ec73d248b974766f8940a896a8c888d8269/lib/levanter/src/levanter/adaptor/lora.py#L158-L162">lora.py</a>).</p> | |
| <p><strong>Merge or distill.</strong> No one has compared the two labs' | |
| post-trained models under one protocol; Marin's panel excludes K2 | |
| Horizon because <a | |
| href="https://github.com/marin-community/marin/issues/9412">"its suite | |
| is incomplete"</a>. The recipes differ on a testable point. IFM trains | |
| separate RL experts, merges their weights, and then repairs the merge | |
| with long SFT or, at 0.9B, distillation. Benjamin Feuer, who leads | |
| Marin's post-training, cautioned that <a | |
| href="https://discord.com/channels/1354881461060243556/1374989195109466122/1551893133514645608">"the | |
| evidence for most of these is pending"</a> before listing, among other | |
| tips, <a | |
| href="https://discord.com/channels/1354881461060243556/1374989195109466122/1551893133514645608">"Mixed-domain | |
| RLVR >> hyperspecialized expert RL from a stability and learning | |
| standpoint <em>but</em> not all regimes benefit equally ... agentic | |
| benefits less"</a> and <a | |
| href="https://discord.com/channels/1354881461060243556/1374989195109466122/1551893133514645608">"RL | |
| after RL shows definite signs of catastrophic forgetting, even mixed | |
| domain"</a>. Marin plans to fold experts into one model by multi-teacher | |
| on-policy distillation (<a | |
| href="https://github.com/marin-community/marin/issues/9250">#9250</a>); | |
| a three-teacher smoke test on Snowball has merged (<a | |
| href="https://github.com/marin-community/MarinSkyRL/pull/749">MarinSkyRL | |
| #749</a>), but Marin has run no weight-merge experiment.</p> | |
| <p><em>Verdict: Marin for RL code, IFM for results. MarinSkyRL is the | |
| more complete and better-tested RL codebase, though it changes weekly | |
| and leans on one engineer; RL360 shows IFM's agentic plumbing but not | |
| the exact code that trained or merged the experts IFM shipped. For SFT | |
| the order flips: IFM's ran on xLLM, whose public code masks | |
| non-assistant tokens, while Marin's Datakit cuts came from an unmerged, | |
| TPU-only branch that trains on every token. Both labs settled on | |
| Megatron for the learner and Harbor for agentic tasks. A team starting | |
| fresh should begin from upstream Miles, which serves rollouts only | |
| through SGLang, or SkyRL, which also supports vLLM, and borrow from | |
| these forks, which are tied to their labs' models and clusters.</em></p> | |
| <h2 id="15-other-factors">15. Other factors</h2> | |
| <ul> | |
| <li><strong>Evaluation.</strong> xLLM ships downstream tasks, including | |
| RULER and SCROLLS, but calls its evaluation <a | |
| href="https://github.com/ifm-ai/xllm/blob/889db388d021c92164e7a675f6884dd40ea117c7/xllm/eval/README.md">"preliminarily | |
| designed for debugging"</a>. Levanter logs per-dataset losses and bits | |
| per byte and integrates lm-eval-harness in process, but Marin scores | |
| post-trained models by serving them on its vLLM fork and running | |
| Evalchemy, lm-eval-harness and Harbor against the endpoint under a | |
| written evaluation policy (<a | |
| href="https://github.com/marin-community/marin/issues/9409">#9409</a>).</li> | |
| <li><strong>Export and serving.</strong> xbridges converts only the | |
| Transformer family, not Gekko; its own vLLM path pins vLLM 0.24.0 and | |
| overwrites the model registry (<a | |
| data-orig-href="https://github.com/ifm-ai/xbridges/blob/227f3fd26e30b024ce09bc7c556c64616c13276e/xbridges/vllm/add_xllm_to_vllm.sh#L4-L20" href="https://github.com/ifm-ai/xbridges/blob/227f3fd26e30b024ce09bc7c556c64616c13276e/xbridges/vllm/add_xllm_to_vllm.sh#L4-L20">script</a>), | |
| though the released models also serve through stock vLLM and SGLang with | |
| remote code. Classic Levanter round-trips about 15 Hugging Face | |
| architectures; Grug exports a custom type served by Marin's vLLM | |
| fork.</li> | |
| <li><strong>Ecosystem and debugging.</strong> PyTorch offers eager | |
| debugging, a larger hiring pool, and the Flash Linear Attention and | |
| Transformer Engine libraries. JAX brings TPUs, but also rack-scale | |
| compiles of about 20 minutes (<a | |
| href="https://github.com/marin-community/marin/issues/8244">#8244</a>) | |
| and compile-time deadlocks, such as ranks auto-tuning different kernel | |
| sizes (<a | |
| href="https://github.com/marin-community/marin/issues/9038">#9038</a>, | |
| since fixed).</li> | |
| <li><strong>People.</strong> xLLM's code has been public for one day, | |
| with one listed contributor; other IFM authors appear in its test paths | |
| and companion repositories. The data toolkit and RL360 went public the | |
| same week, each essentially as one release commit, though RL360's | |
| history survives in IFM's public forks. Marin merged 114 pull requests | |
| in the last full week of September, from 13 people and 4 agent or bot | |
| accounts, and MarinSkyRL merged another 49.</li> | |
| <li><strong>License.</strong> xLLM, the data toolkit, RL360, Marin and | |
| MarinSkyRL use Apache 2.0, though RL360's vendored Megatron-LM keeps | |
| NVIDIA's license. IFM's K2 Horizon datasets use CC-BY-4.0 (TxT360-v2) or | |
| Apache 2.0.</li> | |
| </ul> | |
| <h2 id="recommendation">Recommendation</h2> | |
| <p><strong>For Marin: stay on Levanter.</strong> Switching would abandon | |
| a JAX run 7.96T tokens into 18T, and the scheduler, evaluation and | |
| serving built around it. xLLM lacks two things the hero depends on: | |
| expert parallelism across a rack and Muon-family optimizers. For | |
| frontier MoE, it would add only leading dense layers, dropless routing | |
| and MoVA. Running both trainers would need checkpoint converters and | |
| parity tests that neither side has; Marin's RL work already found the | |
| JAX-to-PyTorch boundary <a | |
| href="https://github.com/marin-community/marin/issues/7164">"orthogonal | |
| to, and harder than, the TPU→GPU (hardware) boundary"</a>. Take four | |
| things from xLLM:</p> | |
| <ol type="1"> | |
| <li><strong>Run the dense benchmark.</strong> Train Llama3-8B at 8K on | |
| the same H100s or H200s in both stacks, scored with one FLOP formula. If | |
| xLLM holds above 43% and Grug trails, fix Grug's Hopper attention | |
| forward before any Hopper-heavy run.</li> | |
| <li><strong>Count causal attention in Levanter's MFU,</strong> or report | |
| both conventions, so numbers compare across stacks and sequence | |
| lengths.</li> | |
| <li><strong>Study IFM's long-context schedule.</strong> It mid-trained | |
| dense models through 32K, 128K and 512K with full attention, the | |
| approach Marin plans for the hero's 65K and 262K phases. The data is | |
| private; the schedule and throughput are public.</li> | |
| <li><strong>Ask IFM about Gekko and MoVA.</strong> They are the two | |
| ideas in xLLM that Marin has not tried, and no released model uses | |
| Gekko. The groups already overlap: Marin serves K2 Horizon models (<a | |
| href="https://github.com/marin-community/marin/pull/9314">#9314</a>), | |
| and an IFM research scientist wrote on Marin's Discord that <a | |
| href="https://discord.com/channels/1354881461060243556/1357057383830126652/1552365242665799831">"the | |
| goals of the organizations seem to be quite aligned"</a>.</li> | |
| </ol> | |
| <p>Four more follow from the data and post-training comparison:</p> | |
| <ol start="5" type="1"> | |
| <li><strong>Test IFM's released data.</strong> The five K2 Horizon | |
| datasets hold about 13.5T tokens, mostly synthetic; TxT360-v2's | |
| <code>web-high-medium</code> subset also carries IFM's quality and topic | |
| labels. Check the ClueWeb licensing, pass the data through Datakit's | |
| deduplication and decontamination, and let a swarm price it like any | |
| other source. Opinion inside Marin is split: Mark Muchane planned to add | |
| <a href="https://github.com/marin-community/marin/issues/8970">"whatever | |
| from the MBZUAI datasets seems actually good"</a>, while David Hall, | |
| after spot-checking one subset, was <a | |
| href="https://discord.com/channels/1354881461060243556/1368297424086499359/1545190602994483270">"kinda | |
| skeptical this will help a large model"</a>.</li> | |
| <li><strong>Favor new data over further mixture search.</strong> New | |
| data gave Marin 1.78× where a curated mix tied proportional sampling, | |
| and IFM reports small gains from a 375,000-GPU-hour search. Grounded | |
| synthetic data is the obvious next source: IFM trained on about 10T | |
| tokens of it, and Marin's own rephrasing research found 1.5–1.8× data | |
| efficiency, yet Datakit has no stage that generates data.</li> | |
| <li><strong>Put the hero's data and SFT code on | |
| <code>main</code>.</strong> Land the quality scorer (<a | |
| href="https://github.com/marin-community/marin/pull/8303">#8303</a>), | |
| restore the topic clustering and mixture search, and merge the release | |
| SFT recipe with a deliberate choice of loss mask, so the next run can | |
| regenerate its data and recipe rather than reload them.</li> | |
| <li><strong>Compare weight merging with distillation.</strong> Merging | |
| is cheap, though IFM follows each merge with long SFT or, at 0.9B, | |
| distillation to repair it. IFM released its 7B experts, merged | |
| checkpoint and merge settings (<a | |
| href="https://wandb.ai/llm360/K2-Horizon-7B/runs/4blppt0e">run</a>), so | |
| Marin can study the method on IFM's checkpoints before training experts | |
| of its own.</li> | |
| </ol> | |
| <p><strong>For other teams,</strong> the hardware decides the | |
| pretraining stack. On TPUs or GB200 racks, choose Levanter. On Hopper | |
| clusters under Slurm, for dense models or long-context research, xLLM is | |
| a credible start, proven to 7B. For frontier MoE on Hopper, xLLM has | |
| probably trained a 375B-A23B model, but with expert parallelism confined | |
| to one node and no published efficiency; compare it with TorchTitan and | |
| Megatron-Core, which already have pipeline parallelism, FP8 and | |
| cross-node expert parallelism. If you pretrain in JAX, budget for a | |
| PyTorch port of each model for RL: Marin's first port ran 27× slower | |
| until grouped kernels brought it to 1.4×, and its parity checks stop at | |
| tiny models. For data, Datakit is the more complete and better-tested | |
| pipeline, though beyond one machine it needs Marin's Iris scheduler; | |
| from IFM, the released datasets are worth more than the toolkit. For RL, | |
| start from upstream Miles or SkyRL and borrow from the forks: IFM's | |
| Miles fork for Harbor-based agentic rollouts, and MarinSkyRL for router | |
| replay, per-expert weight sync and its Megatron port of a JAX-trained | |
| MoE.</p> | |
| <h2 id="method-and-limits">Method and limits</h2> | |
| <p>Thirty-one research agents read the seven main repositories and IFM's | |
| related data repositories, searched Marin's records and gathered IFM's | |
| public materials; six more reviewed the drafts adversarially, and I | |
| checked the claims behind each verdict against code, logs and threads | |
| myself. A follow-up pass examined the 375B's Hugging Face files and | |
| W&B project, and I read IFM's data talk slide by slide. Nothing ran | |
| on accelerators; one agent ran the data toolkit's CPU stages on | |
| synthetic data. IFM assembled its public W&B projects after the fact | |
| from private runs, so its GPU counts and hardware are inferred, and IFM | |
| runs internal code beyond its public releases. My token counts for IFM's | |
| datasets extrapolate from sampled rows and could be off by a fifth or | |
| more. MarinSkyRL's own issues are outside Marin's mirrored records, so I | |
| read them directly on GitHub.</p> | |
| </div> | |
| <footer class="provenance"> | |
| <p><em>Corpus: 2026-09-28 17:07 UTC · Generated: 2026-09-29 08:01 UTC · Viewing: <span id="viewing-time"></span></em></p> | |
| <blockquote> | |
| <p><em>Data: marinmirror — 225,653 chunks, built 2026-09-28 17:07 UTC · | |
| summaries through 2026-09-21_2026-09-27. Code: xLLM 889db38, xattn | |
| 0d6d73b, xbridges 227f3fd, pretraining-data-toolkit cd93e11, RL360 | |
| 5b548b6, search360 4f9a64d, TxT360 07d98df, LCQA 1dee244, | |
| PRism-synthesis 7fd0264, process_entry a4745b1, Marin main 4aa26ec | |
| (release-SFT branch 8a0ec6c), MarinSkyRL 0cdccc9. IFM sources: launch | |
| blog and product page (Wayback), Hugging Face model and dataset cards | |
| and dataset statistics (375B at revision 82af3bc), public W&B | |
| projects (entities llm360 and mbzuai-llm), the Diversity First data | |
| talk, arXiv 2601.06463, 2512.06201 and 2510.24397; TorchTitan benchmark | |
| docs. Marin sources also include the Datakit blog post (openathena.ai) | |
| and Marin's Hugging Face artifacts.</em></p> | |
| <p><em>Query: "Do a detailed analysis of <a | |
| href="https://github.com/ifm-ai/xllm">https://github.com/ifm-ai/xllm</a>, | |
| a new library for specifying and training large LLMs, and compare it to | |
| what we have in Levanter. Consider flexibility in specifying new | |
| architectures, completeness of kinds of linear and other scalable | |
| attention mechanisms implemented, FP8 and FP4 training, completeness of | |
| kernels for different GPUs, ability to train on non-GPUs, dimensions of | |
| parallelism implemented, ease of evolution by agents, code complexity, | |
| test coverage, performance, and any other measure important for a team | |
| building frontier LLMs who needs to choose between Levanter and xLLM. | |
| Write a detailed report, then do an adversarial review, then publish a | |
| gist." Follow-ups: "Does xLLM support sparse MoE architectures? Is there | |
| evidence they used this codebase to train their largest model at <a | |
| href="https://huggingface.co/IFM/K2-Horizon-375B-A23B">https://huggingface.co/IFM/K2-Horizon-375B-A23B</a>?" | |
| and "Use the same protocol of detailed analysis and adversarial review | |
| to update the report with data curation and mixing code from <a | |
| href="https://github.com/ifm-ai/pretraining-data-toolkit">https://github.com/ifm-ai/pretraining-data-toolkit</a> | |
| versus datakit from Marin as well as <a | |
| href="https://github.com/ifm-ai/RL360">https://github.com/ifm-ai/RL360</a> | |
| versus MarinSkyRL and other post-training code from Levanter and Marin. | |
| Update the gist when you finish. Include information for IFM about data | |
| curation and mixing from the presentation at <a | |
| href="https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf">https://moonfolk.github.io/files/data-diversity-april-9-talk.pdf</a>."</em></p> | |
| <p><em>Sub-queries: xLLM architecture and config system · xLLM attention | |
| and sequence mixers (xattn, FLA) · xLLM kernels, precision and hardware | |
| · xLLM parallelism and checkpointing · xLLM tests, quality and benchmark | |
| claims · xLLM external context (IFM, K2 Horizon, W&B logs) · | |
| Levanter/Grug architecture flexibility · Levanter attention and linear | |
| mixers · Levanter kernels, FP8 and accelerators · Levanter parallelism | |
| and fault tolerance · Levanter tests, docs and agent affordances · Marin | |
| measured MFU · Marin FP8/FP4 history · Marin attention and hybrid-mixer | |
| experiments · Marin parallelism decisions · Marin kernel engineering and | |
| maintenance · Marin TPU and non-NVIDIA use · Marin agent-driven | |
| development · Marin JAX-vs-PyTorch design decisions · Marin long-context | |
| work and mentions of xLLM/IFM · in-flight code branches · adversarial | |
| reviews: xLLM fact-check, Levanter fact-check, | |
| prose/completeness/fairness · follow-up: xLLM MoE support; | |
| K2-Horizon-375B-A23B provenance (card, config, model code, W&B runs, | |
| run ID, Megatron option names) · round 2: IFM pretraining-data-toolkit | |
| code · IFM data practice (talk, blog, released datasets, K2-V2 paper) · | |
| Marin Datakit code · Marin Datakit history and hero data · Marin | |
| data-mixing methods and results · IFM RL360 code and W&B runs · | |
| MarinSkyRL code · Marin RL infrastructure history · Marin post-training | |
| recipes and results · Levanter post-training code · adversarial reviews | |
| (round 2): IFM-side fact-check, Marin-side fact-check, prose, | |
| completeness and fairness.</em></p> | |
| </blockquote> | |
| </footer> | |
| </div> | |
| </body> | |
| </html> |
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment