Evaluated with GPT-4o Judge across 6 core Korean language & reasoning domains.
| Model | Single Turn | Multi Turn | Overall Score (10점 만점) |
|---|---|---|---|
Gaiel-1.5B-Korean-Tuned-MLX |
2.14 | 1.59 | 1.87 |
Gaiel-7B-Korean-Tuned-MLX |
4.75 | 3.60 | 4.18 |
{
"runs": [
{
"elapsed_s": 4.551236792001873,
"output_text_preview": "A: \"a\" is a single character with a count of one. 0 characters are needed to represent this. \n\nB: \"a\" is the same as \"a a\". 0 characters are needed to represent this. \n\nC: \"abc\" is a three-character s",Status: Model exceeds currently available cluster RAM limit for unquantized inference (Requires additional unified memory or 4-bit quantization).
{
"runs": [
{
"elapsed_s": 4.551236792001873,
"output_text_preview": "A: \"a\" is a single character with a count of one. 0 characters are needed to represent this. \n\nB: \"a\" is the same as \"a a\". 0 characters are needed to represent this. \n\nC: \"abc\" is a three-character s",
"stats": {Failed: Command failed with exit code 143
Failed: Command failed with exit code -15
Failed: Command failed with exit code 1
Failed: Command failed with exit code 1