Skip to content

Commit 7717121

Browse files
committed
Align performance docs with committed benchmark results
1 parent 0a85951 commit 7717121

6 files changed

Lines changed: 115 additions & 18 deletions

File tree

‎README.md‎

Lines changed: 5 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@
44

55
**Open-weight typed decisions, running natively on Apple Silicon.**
66

7-
**13.4 ms** median end-to-end for a short English typed decision. **7.4 ms** with the multilingual checkpoint. **0 output tokens.** Local MLX inference, with no PyTorch, Transformers runtime, or cloud API.
7+
**17.75 ms** median end-to-end for a short English typed decision. **10.91 ms** with the multilingual checkpoint. **0 output tokens.** Local MLX inference, with no PyTorch, Transformers runtime, or cloud API.
88

99
[中文](https://github.com/mizorewww/laya-mlx/blob/main/README.zh-CN.md) · [Benchmarks](https://github.com/mizorewww/laya-mlx/blob/main/BENCHMARKS.md) · [Snake demo](https://github.com/mizorewww/laya-mlx/blob/main/docs/SNAKE_DEMO.md) · [Hugging Face weights](https://huggingface.co/aac6fef/laya-mlx)
1010

@@ -51,12 +51,12 @@ Download once before the offline demo. Use a terminal at least 104 × 35 cells.
5151

5252
| FP16, end-to-end | Laya 421M | Multilingual 322M |
5353
|---|---:|---:|
54-
| One short question, P50 | **13.42 ms** | **7.39 ms** |
55-
| One short question, P95 | **13.92 ms** | **7.79 ms** |
56-
| 50-question throughput | **146.8 q/s** | **395.0 q/s** |
54+
| One short question, P50 | **17.75 ms** | **10.91 ms** |
55+
| One short question, P95 | **21.45 ms** | **19.48 ms** |
56+
| 50-question throughput | **143.3 q/s** | **402.2 q/s** |
5757
| Peak MLX allocation, one short question | **943.6 MiB** | **687.6 MiB** |
5858

59-
M3 Max, 40 GPU cores, 128 GiB memory. Timing includes prompt preparation, tokenization, tensors, synchronized inference, calibration and result formatting; model loading is excluded. The 50-question measurement uses `batch_size=64`; the API defaults to 16. Different lengths, question counts and runtime conditions change latency. [Full method and every timing sample](https://github.com/mizorewww/laya-mlx/blob/main/BENCHMARKS.md).
59+
M3 Max, 40 GPU cores, 128 GiB memory. These figures use the committed 2026-09-22 benchmark run. Timing includes prompt preparation, tokenization, tensors, synchronized inference, calibration and result formatting; model loading is excluded. The 50-question measurement uses `batch_size=64`; the API defaults to 16. Different lengths, question counts and runtime conditions change latency. [Full method and every timing sample](https://github.com/mizorewww/laya-mlx/blob/main/BENCHMARKS.md).
6060

6161
**Port fidelity:** all three checkpoints matched the upstream selected answer on **63/63 validation questions in both FP32 and FP16** — 378/378 comparisons. Each configuration also passed 100 repeated finite, deterministic calls with zero measured active-memory growth. This measures fidelity on those fixtures, not accuracy on every possible question. [Probability errors and validation](https://github.com/mizorewww/laya-mlx/blob/main/BENCHMARKS.md#numerical-parity-and-stability).
6262

‎README.zh-CN.md‎

Lines changed: 6 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -4,11 +4,11 @@
44

55
**在 Apple Silicon 上本地运行开放权重的结构化决策模型。**
66

7-
单个短问题端到端中位耗时 **13.4 ms**;multilingual 检查点为 **7.4 ms**。**0 个输出 token**,原生 MLX,无 PyTorch / Transformers 推理依赖,无云端 API。
7+
单个短问题端到端中位耗时 **17.75 ms**;multilingual 检查点为 **10.91 ms**。**0 个输出 token**,原生 MLX,无 PyTorch / Transformers 推理依赖,无云端 API。
88

99
[English](README.md) · [完整 benchmark](BENCHMARKS.md) · [Snake 使用说明](docs/SNAKE_DEMO.md) · [30 秒 MP4](docs/assets/snake-demo.mp4)
1010

11-
GIF 使用真实游戏记录按原始时间戳渲染。每步都调用 Laya,界面显示循环路径安全层及其接管次数。上面的 13.4 / 7.4 ms 来自**单问题 API 基准**,并非每步批量回答三个问题的 Snake 帧耗时;游戏速度见[独立报告](docs/SNAKE_BENCHMARKS.md)。
11+
GIF 使用真实游戏记录按原始时间戳渲染。每步都调用 Laya,界面显示循环路径安全层及其接管次数。上面的 17.75 / 10.91 ms 来自**单问题 API 基准**,并非每步批量回答三个问题的 Snake 帧耗时;游戏速度见[独立报告](docs/SNAKE_BENCHMARKS.md)。
1212

1313
## 快速开始
1414

@@ -49,12 +49,12 @@ laya-snake
4949

5050
| FP16,端到端 | Laya 421M | Multilingual 322M |
5151
|---|---:|---:|
52-
| 单个短问题 P50 | **13.42 ms** | **7.39 ms** |
53-
| 单个短问题 P95 | **13.92 ms** | **7.79 ms** |
54-
| 50 问题吞吐量 | **146.8 q/s** | **395.0 q/s** |
52+
| 单个短问题 P50 | **17.75 ms** | **10.91 ms** |
53+
| 单个短问题 P95 | **21.45 ms** | **19.48 ms** |
54+
| 50 问题吞吐量 | **143.3 q/s** | **402.2 q/s** |
5555
| 单个短问题 MLX 峰值分配 | **943.6 MiB** | **687.6 MiB** |
5656

57-
硬件为 M3 Max(40 核 GPU、128 GiB 内存)。计时包含提示准备、tokenization、张量构建、GPU 同步推理、校准及结果格式化,排除模型加载。50 问题吞吐量使用 `batch_size=64`,公开 API 默认为 16。
57+
硬件为 M3 Max(40 核 GPU、128 GiB 内存)。数据来自仓库中 2026-09-22 的基准测试。计时包含提示准备、tokenization、张量构建、GPU 同步推理、校准及结果格式化,排除模型加载。50 问题吞吐量使用 `batch_size=64`,公开 API 默认为 16。
5858

5959
**移植一致性:**三个检查点在 FP32 和 FP16 下均通过 **63/63** 验证问题的上游 argmax 对齐,合计 378/378;每个配置各执行 100 次重复调用,结果有限、确定,测得活跃内存增长为零。它验证移植保真度,不代表所有实际问题都能答对。[完整误差和原始记录](BENCHMARKS.md)。
6060

‎docs/LAUNCH.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -21,7 +21,7 @@ The extra maximum-speed clip comes from a separate 20.01-second truecolor TTY ru
2121
>
2222
> Laya-MLX runs open-weight typed decision models on Apple Silicon. Watch a 322M model play Snake with a visible cycle safety layer: real probabilities, measured latency, 0 output tokens, no inference API.
2323
>
24-
> One-question API benchmark: 7.39 ms p50 on M3 Max.
24+
> One-question API benchmark: 10.91 ms p50 on M3 Max.
2525
>
2626
> `pip install laya-mlx`
2727
>
@@ -33,7 +33,7 @@ The extra maximum-speed clip comes from a separate 20.01-second truecolor TTY ru
3333
>
3434
> Laya-MLX:在 Mac 上本地运行的开放权重决策模型。这个 3.22 亿参数的贪吃蛇 demo,每一步都显示真实方向概率、推理耗时和安全层接管次数。
3535
>
36-
> 0 个输出 token,无推理 API。M3 Max 单问题基准 P50 为 7.39 ms。
36+
> 0 个输出 token,无推理 API。M3 Max 单问题基准 P50 为 10.91 ms。
3737
>
3838
> `pip install laya-mlx`
3939
>
@@ -43,4 +43,4 @@ The extra maximum-speed clip comes from a separate 20.01-second truecolor TTY ru
4343

4444
The optimized complete Snake loop measured **75.40 moves/second over 2,400 moves**, with zero deaths, 2 safety interventions and 2,400/2,400 executed-action agreement with the paired eager control. It was about **6.5% faster in that run**. This includes planning, inference, Rich composition, ANSI serialization and game updates, but excludes the terminal emulator's painting.
4545

46-
Use the [optimization report](SNAKE_OPTIMIZATION.md) when sharing that number. The 7.39 ms headline describes the separate one-question API fixture; it is not the frame time of this three-question Snake demonstration. Neither result is a cloud-API comparison or evidence of unaided Snake reasoning.
46+
Use the [optimization report](SNAKE_OPTIMIZATION.md) when sharing that number. The 10.91 ms headline describes the separate one-question API fixture from the committed 2026-09-22 run; it is not the frame time of this three-question Snake demonstration. Neither result is a cloud-API comparison or evidence of unaided Snake reasoning.

‎docs/MATH_10X_RESEARCH.md‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,6 @@
11
# Can Laya MLX become ten times faster? A mathematical investigation
22

3-
Research date: 2026-09-19. Baseline: the committed FP16 MLX results on the Apple M3 Max with 40 GPU cores and 128 GiB unified memory. This report separates **algebraic facts**, **static cost estimates**, **CPU measurements of selected checkpoint matrices**, and **hypotheses requiring inference experiments**. No GPU inference or new latency measurement was performed for this mathematical investigation. The companion [engineering investigation](ENGINEERING_10X_RESEARCH.md) contains candidate timings when available. Here, “exact” refers to preserving mathematical dependencies and the real-arithmetic function; a different GPU reduction order or kernel can still change floating-point results, so the existing numerical tolerances and exposed-output contract remain acceptance gates.
3+
Research date: 2026-09-19. Baseline: the committed FP16 MLX results at `49daed6` on the Apple M3 Max with 40 GPU cores and 128 GiB unified memory. This report separates **algebraic facts**, **static cost estimates**, **CPU measurements of selected checkpoint matrices**, and **hypotheses requiring inference experiments**. No GPU inference or new latency measurement was performed for this mathematical investigation. The companion [engineering investigation](ENGINEERING_10X_RESEARCH.md) contains candidate timings when available. Here, “exact” refers to preserving mathematical dependencies and the real-arithmetic function; a different GPU reduction order or kernel can still change floating-point results, so the existing numerical tolerances and exposed-output contract remain acceptance gates.
44

55
**Decision:** do not budget for a universal 10× end-to-end improvement from hand-written kernels while retaining these checkpoints and their full outputs. Exact local attention and output pruning are worthwhile, bounded improvements. Direct low-rank decomposition is not close to lossless in the four sampled weight matrices. A 10× product improvement is credible for workloads with substantial exact repetition, or as the goal of a substantially smaller distilled/restructured model. Those are different promises and must have different benchmarks.
66

@@ -14,7 +14,7 @@ These are the existing synchronized, warm end-to-end medians, including preparat
1414
| Multilingual | 7.390 → **0.739 ms** | 27.386 → 2.739 ms | 127.565 → 12.756 ms | 37.635 → **3.763 ms** | 389.487 → 38.949 ms |
1515
| Typed decisions | 13.712 → **1.371 ms** | 75.618 → 7.562 ms | 380.560 → 38.056 ms | 99.233 → **9.923 ms** | 1000.294 → 100.029 ms |
1616

17-
Sources: [Laya FP16](../benchmarks/results/laya-mlx-float16.json), [multilingual FP16](../benchmarks/results/laya-multilingual-mlx-float16.json), [typed-decisions FP16](../benchmarks/results/laya-typed-decisions-mlx-float16.json). Short padded lengths are 93/91/93; long lengths are 512/1024/1024. Comparing their long rows does not hold token length constant. The target is a further improvement over native MLX FP16, not over PyTorch MPS FP32.
17+
Sources at the research snapshot: [Laya FP16](https://github.com/mizorewww/laya-mlx/blob/49daed6609cb3da142d1a0c88e538dc07f00d974/benchmarks/results/laya-mlx-float16.json), [multilingual FP16](https://github.com/mizorewww/laya-mlx/blob/49daed6609cb3da142d1a0c88e538dc07f00d974/benchmarks/results/laya-multilingual-mlx-float16.json), [typed-decisions FP16](https://github.com/mizorewww/laya-mlx/blob/49daed6609cb3da142d1a0c88e538dc07f00d974/benchmarks/results/laya-typed-decisions-mlx-float16.json). Short padded lengths are 93/91/93; long lengths are 512/1024/1024. Comparing their long rows does not hold token length constant. The target is a further improvement over native MLX FP16, not over PyTorch MPS FP32.
1818

1919
For any proposed optimization, let `f` be its **measured fraction of end-to-end wall time** and `s` its own acceleration. Amdahl's law gives:
2020

‎docs/PERFORMANCE_RESEARCH.md‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,6 @@
11
# Laya MLX performance research
22

3-
Research date: 2026-09-19. Target: Apple M3 Max, 40 GPU cores, 128 GiB unified memory, MLX/MLX Metal 0.32.2. This is a static review of the native runtime, installed MLX implementation, official documentation, and existing benchmark JSON. **No GPU benchmark or model inference was run for this research. None of the proposed optimizations below has a measured speedup in this report.**
3+
Research date: 2026-09-19. Benchmark snapshot: commit `49daed6` from that date. Target: Apple M3 Max, 40 GPU cores, 128 GiB unified memory, MLX/MLX Metal 0.32.2. This is a static review of the native runtime, installed MLX implementation, official documentation, and existing benchmark JSON. **No GPU benchmark or model inference was run for this research. None of the proposed optimizations below has a measured speedup in this report.**
44

55
The first experiments should be whole-model compilation and representative batch scheduling, followed by selective quantized matrix multiplication. These address the dominant repeated work. A specialized local-attention kernel is a credible longer-term project for long inputs. Exact pruning of the last decision-head layer is feasible, but its whole-model arithmetic saving is only a few percent. Large improvements without changing the checkpoint will require improving the dense backbone, eliminating genuinely redundant requests, or finding a measured implementation bottleneck; simply replacing an activation or enabling another attention flag is unlikely to suffice.
66

@@ -17,7 +17,7 @@ The following are existing **end-to-end median latencies**, including prompt pre
1717
| Multilingual MLX FP32 | 7.988 ms | 32.337 ms | 151.387 ms | 47.208 ms | 451.331 ms |
1818
| Multilingual stock Torch MPS FP32 | 19.349 ms | 43.158 ms | 194.171 ms | 52.939 ms | 534.492 ms |
1919

20-
Sources: [Laya FP16](../benchmarks/results/laya-mlx-float16.json), [Laya FP32](../benchmarks/results/laya-mlx-float32.json), [Laya MPS](../benchmarks/results/laya-torch-mps-float32.json), [multilingual FP16](../benchmarks/results/laya-multilingual-mlx-float16.json), [multilingual FP32](../benchmarks/results/laya-multilingual-mlx-float32.json), and [multilingual MPS](../benchmarks/results/laya-multilingual-torch-mps-float32.json). Long inputs contain 512 tokens for Laya and 1024 for multilingual; comparing their long-input latencies is therefore not a comparison at equal sequence length. Short padded lengths are 93 and 91, respectively. Torch comparisons must retain the FP32 label: they combine a backend change with a precision change when compared against MLX FP16.
20+
Sources at the research snapshot: [Laya FP16](https://github.com/mizorewww/laya-mlx/blob/49daed6609cb3da142d1a0c88e538dc07f00d974/benchmarks/results/laya-mlx-float16.json), [Laya FP32](https://github.com/mizorewww/laya-mlx/blob/49daed6609cb3da142d1a0c88e538dc07f00d974/benchmarks/results/laya-mlx-float32.json), [Laya MPS](https://github.com/mizorewww/laya-mlx/blob/49daed6609cb3da142d1a0c88e538dc07f00d974/benchmarks/results/laya-torch-mps-float32.json), [multilingual FP16](https://github.com/mizorewww/laya-mlx/blob/49daed6609cb3da142d1a0c88e538dc07f00d974/benchmarks/results/laya-multilingual-mlx-float16.json), [multilingual FP32](https://github.com/mizorewww/laya-mlx/blob/49daed6609cb3da142d1a0c88e538dc07f00d974/benchmarks/results/laya-multilingual-mlx-float32.json), and [multilingual MPS](https://github.com/mizorewww/laya-mlx/blob/49daed6609cb3da142d1a0c88e538dc07f00d974/benchmarks/results/laya-multilingual-torch-mps-float32.json). Long inputs contain 512 tokens for Laya and 1024 for multilingual; comparing their long-input latencies is therefore not a comparison at equal sequence length. Short padded lengths are 93 and 91, respectively. Torch comparisons must retain the FP32 label: they combine a backend change with a precision change when compared against MLX FP16.
2121

2222
There is meaningful run variability. For example, the multilingual FP16 long 10-question run has p50 389.487 ms, p95 462.319 ms, and maximum 619.663 ms. Its short single-question forward median is 8.023 ms while the independently measured end-to-end median is 7.390 ms. Subtracting those medians would produce a nonsensical negative preprocessing time. The current files do **not** isolate tokenizer, Python dispatch, individual GPU kernels, or synchronization costs. They establish useful baselines, not a kernel-level bottleneck diagnosis.
2323

‎tests/test_benchmark_docs.py‎

Lines changed: 97 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,97 @@
1+
"""Keep current-facing performance figures aligned with the committed results."""
2+
3+
import json
4+
from pathlib import Path
5+
6+
ROOT = Path(__file__).resolve().parents[1]
7+
RESULTS = ROOT / "benchmarks" / "results"
8+
HISTORICAL_REVISION = "49daed6609cb3da142d1a0c88e538dc07f00d974"
9+
10+
11+
def metric(model: str, questions: int, key: str) -> float:
12+
report = json.loads((RESULTS / f"{model}-mlx-float16.json").read_text())
13+
row = next(
14+
result
15+
for result in report["results"]
16+
if result["workload"] == "short" and result["questions"] == questions
17+
)
18+
return row["end_to_end"][key]
19+
20+
21+
def test_current_readmes_match_committed_benchmarks() -> None:
22+
english = ROOT.joinpath("README.md").read_text()
23+
chinese = ROOT.joinpath("README.zh-CN.md").read_text()
24+
laya = "laya"
25+
multilingual = "laya-multilingual"
26+
laya_p50 = f"{metric(laya, 1, 'p50_ms'):.2f}"
27+
multilingual_p50 = f"{metric(multilingual, 1, 'p50_ms'):.2f}"
28+
run_date = json.loads((RESULTS / "laya-mlx-float16.json").read_text())["created_at"].split("T")[
29+
0
30+
]
31+
32+
assert (
33+
f"**{laya_p50} ms** median end-to-end for a short English typed decision. "
34+
f"**{multilingual_p50} ms** with the multilingual checkpoint."
35+
in "\n".join(english.splitlines()[:8])
36+
)
37+
assert (
38+
f"单个短问题端到端中位耗时 **{laya_p50} ms**;multilingual 检查点为 "
39+
f"**{multilingual_p50} ms**。" in "\n".join(chinese.splitlines()[:8])
40+
)
41+
assert f"上面的 {laya_p50} / {multilingual_p50} ms 来自" in chinese
42+
assert f"committed {run_date} benchmark run" in english
43+
assert f"{run_date} 的基准测试" in chinese
44+
45+
for text, labels in (
46+
(english, ("One short question, P50", "One short question, P95", "50-question throughput")),
47+
(chinese, ("单个短问题 P50", "单个短问题 P95", "50 问题吞吐量")),
48+
):
49+
for label, questions, key, precision, unit in (
50+
(labels[0], 1, "p50_ms", 2, "ms"),
51+
(labels[1], 1, "p95_ms", 2, "ms"),
52+
(labels[2], 50, "questions_per_second", 1, "q/s"),
53+
):
54+
expected = (
55+
f"| {label} | **{metric(laya, questions, key):.{precision}f} {unit}**"
56+
f" | **{metric(multilingual, questions, key):.{precision}f} {unit}** |"
57+
)
58+
assert expected in text
59+
60+
61+
def test_release_copy_matches_committed_multilingual_latency() -> None:
62+
text = ROOT.joinpath("docs", "LAUNCH.md").read_text()
63+
latency = f"{metric('laya-multilingual', 1, 'p50_ms'):.2f} ms"
64+
run_date = json.loads((RESULTS / "laya-multilingual-mlx-float16.json").read_text())[
65+
"created_at"
66+
].split("T")[0]
67+
assert f"One-question API benchmark: {latency} p50" in text
68+
assert f"单问题基准 P50 为 {latency}" in text
69+
assert f"The {latency} headline describes" in text
70+
assert f"committed {run_date} run" in text
71+
72+
73+
def test_historical_research_links_keep_the_original_results() -> None:
74+
reports = {
75+
"PERFORMANCE_RESEARCH.md": (
76+
"laya-mlx-float16",
77+
"laya-mlx-float32",
78+
"laya-torch-mps-float32",
79+
"laya-multilingual-mlx-float16",
80+
"laya-multilingual-mlx-float32",
81+
"laya-multilingual-torch-mps-float32",
82+
),
83+
"MATH_10X_RESEARCH.md": (
84+
"laya-mlx-float16",
85+
"laya-multilingual-mlx-float16",
86+
"laya-typed-decisions-mlx-float16",
87+
),
88+
}
89+
for report, sources in reports.items():
90+
text = ROOT.joinpath("docs", report).read_text()
91+
assert "`49daed6`" in text
92+
for source in sources:
93+
assert (
94+
"https://github.com/mizorewww/laya-mlx/blob/"
95+
f"{HISTORICAL_REVISION}/benchmarks/results/{source}.json"
96+
) in text
97+
assert f"(../benchmarks/results/{source}.json)" not in text

0 commit comments

Comments
 (0)