This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| # The most atomic way to train and inference a GPT in pure, dependency-free Bend. | |
| # This file is the complete algorithm. | |
| # Everything else is just efficiency. | |
| # | |
| # Port of microgpt.py by @karpathy. | |
| # | |
| # Bend is pure and affine, so there is no mutable object graph. A Value is a | |
| # pair V{id, data}; every op appends one node (its children ids and local | |
| # grads) to a tape threaded by a state monad G. Since ids grow with creation | |
| # order, the tape read newest-first is already the reversed topological order |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| # microgpt_atomics_inline.exs | |
| # Port of microgpt.py by @karpathy | |
| # | |
| # Storage: :atomics array (7 slots per node) in place of Python dicts. | |
| # Refs threaded: both data ref and counter ref passed as args, zero | |
| # :persistent_term.get calls during training/inference. | |
| # @compile inline on all hot V and MicroGPT functions. | |
| # | |
| # Optimizations that diverge from the Python original: | |
| # |