Ael
Ael Benchmarks
Audited on NVIDIA A100-SXM4-40GB • 2.15M to 2.77M Real Tokens • Zero Synthetic Padding

Multi-Million-Token Intelligence. Bounded in 2,112 Active Tokens.

Conventional Transformers require 57.3 GiB to 137.3 GiB of Key-Value cache memory alone at 2.15M–2.77M tokens. Powered by the ISOM-R2 Paged Virtual SVD Cache and $SO(d)$ Lie-manifold transport, the four models in the Ael Model Family stream and synthesize across entire open-source repositories in 7.42 to 13.94 seconds with +0.01 GB to +0.09 GB streaming KV overhead.

01 • Model Lineup

Four Specialized Architectures. One Bounded State Engine.

Active GPU Attention Window: 2,112 Tokens (64 Sinks + 2,048 Window)
1.54B Dense 28L GQA

Ael-Coder-1.5B

Multi-file PyTorch code synthesis across 82 transformers modules. Synthesizes LoggedGELU across a 1.08M-token inter-file gap with zero numerical error.

Context
2,147,447
Throughput
289.5K/s
Peak VRAM
3.14 GB
Total Time
7.42 s
Open Model Card →
1.54B Dense SO(64) Lie Verifier

Ael-Reasoning-1.5B

Symbolic mathematics & combinatorics across 42 sympy modules. Synthesizes exact DerangedFibonacci recurrences with $O(1)$ Lie-manifold rollback.

Context
2,311,513
Throughput
231.7K/s
Peak VRAM
3.03 GB
Total Time
9.97 s
Open Model Card →
15.71B / 2.36B Act 64-Expert MLA

Ael-Coder-16B-MoE

Repository-scale Mixture-of-Experts with Multi-Head Latent Attention ($SO(192)$). Bridges a 1.39M-token inter-file gap with +0.06 GB streaming KV overhead.

Context
2,774,027
Throughput
227.7K/s
Peak VRAM
30.51 GB
Total Time
12.18 s
Open Model Card →
40.0B Dense 60L 4-Bit NF4

Ael-Pro-40B

Enterprise backend security & cryptographic compliance across 122 django modules. Synthesizes 1.8M-iteration PBKDF2 + HMAC-SHA256 audit classes.

Context
2,399,330
Throughput
172.1K/s
Peak VRAM
24.73 GB
Total Time
13.94 s
Open Model Card →

02 • Hardware Telemetry

Memory Compression & Streaming Speed

Measured on single NVIDIA A100-SXM4-40GB GPU

Total GPU Memory at 2.15M–2.77M Tokens

Ael Audited Peak VRAM (Weights + 2,112 Active KV) vs. Standard Transformer (Weights + Full FP16 KV)

Lower is Better
Ael-Coder-1.5B (2.15M tokens) 19.2× Lower VRAM
Ael (ISOM-R2)
3.14 GB
Standard GQA
60.32 GB
Ael-Reasoning-1.5B-Instruct (2.31M tokens) 21.3× Lower VRAM
Ael (ISOM-R2)
3.03 GB
Standard GQA
64.61 GB
Ael-Coder-16B-MoE (2.77M tokens) +0.06 GB Stream KV
Ael (ISOM-R2)
30.51 GB
Standard MLA
115.05 GB
Ael-Pro-40B (2.40M tokens) 6.4× Lower VRAM
Ael (ISOM-R2)
24.73 GB
Standard MQA
158.91 GB

Audited Telemetry Chart

Switch between Streaming Throughput (tok/s) and Memory Scaling (GB)

Zero OOM Guarantee: Every Ael model completes 2.15M to 2.77M real-token prefill and multi-hop generation well inside a single 40GB A100 GPU, sustaining 172,072 to 289,500 tokens/sec.

03 • Benchmark Matrix

Complete Audited Evaluation Table

Model & Class Target Domain & Real Corpus Audited Tokens Inter-Hop Gap Time & Speed Weights / Peak VRAM Multi-Hop Recall & Live Verification
Ael-Coder-1.5B
AelCoder15BForCausalLM • 1.54B
Multi-File PyTorch Synthesis
huggingface/transformers (82 files)
2,147,447
1,049 chunks
1,077,049 tok
526 chunks
7.42 s
289,500 tok/s
2.98 GB / 3.14 GB
+0.01 GB stream
Pages [388, 791, 265] (100%)
Synthesized LoggedGELU • max_err = 0.00e+00
Ael-Reasoning-1.5B-Instruct
AelReasoning15BForCausalLM • 1.54B
Symbolic Math & Combinatorics
sympy/sympy (42 modules)
2,311,513
1,129 chunks
932,802 tok
455 chunks
9.97 s
231,748 tok/s
2.88 GB / 3.03 GB
+0.01 GB stream
Pages [501, 956] (100%)
Synthesized DerangedFibonacci • $SO(64)$ 2.51e-05
Ael-Coder-16B-MoE
AelCoder16BMoEForCausalLM • 15.71B/2.36B
Repository-Scale MoE Synthesis
huggingface/transformers (82 files)
2,774,027
1,355 chunks
1,394,629 tok
681 chunks
12.18 s
227,686 tok/s
29.28 GB / 30.51 GB
+0.06 GB stream
Pages [339, 1020] (100%)
Synthesized LoggedGELU • $SO(192)$ MLA 4.47e-05
Ael-Pro-40B
AelPro40BForCausalLM • 40.0B NF4
Enterprise Security & Crypto Audit
django/django (122 modules)
2,399,330
1,172 chunks
1,171,652 tok
572 chunks
13.94 s
172,072 tok/s
21.62 GB / 24.73 GB
+0.09 GB stream
Pages [304, 876] (100%)
Synthesized SignedPBKDF2Hasher • HMAC + Tamper Pass

04 • Audited Execution Logs

Verbatim A100 Hardware Transcripts

Select any of the four benchmarked Ael models to inspect its streaming prefill telemetry, retrieved pages, and live execution verification.

Ael-Coder-1.5B — 2,147,447-Token Multi-File Code Synthesis

82 Python files from huggingface/transformers • Hop 1: CaptureStdout (Chunk 265) • Hop 2: GELUActivation (Chunk 791)

nvidia-a100-sxm4-40gb — Prannesshkva/Ael-Coder-1.5B
Loaded AelCoder15BForCausalLM | Model Weights VRAM: 2.98 GB
Tokenizing 2.15M+ real-world multi-file corpus (82 files | 9,780,715 chars)...
Hop 1 Target   : class CaptureStdout(CaptureStd) @ Token #543,865 -> Chunk 265
Hop 2 Target   : class GELUActivation(nn.Module) @ Token #1,620,914 -> Chunk 791
Inter-File Gap : 1,077,049 real tokens (526 chunks apart)

  [ISOM-R2 Engine] Streaming 2,147,383 codebase context tokens across 1049 chunks (active GPU buffer < 400 MB)...
  [ISOM-R2] Prefill 428,032 / 2,147,383 tokens (19.9%) | Active buffer: 2112 tokens | VRAM: 2.99 GB
  [ISOM-R2] Prefill 856,064 / 2,147,383 tokens (39.9%) | Active buffer: 2112 tokens | VRAM: 2.99 GB
  [ISOM-R2] Prefill 1,284,096 / 2,147,383 tokens (59.8%) | Active buffer: 2112 tokens | VRAM: 2.99 GB
  [ISOM-R2] Prefill 1,712,128 / 2,147,383 tokens (79.7%) | Active buffer: 2112 tokens | VRAM: 2.99 GB
  [ISOM-R2] Prefill 2,140,160 / 2,147,383 tokens (99.7%) | Active buffer: 2112 tokens | VRAM: 2.99 GB
  [ISOM-R2] Prefill 2,147,383 / 2,147,383 tokens (100.0%) | Active buffer: 2112 tokens | VRAM: 2.99 GB
  [ISOM-R2] Retrieved salient context pages: [388, 791, 265] | Active KV: 2112 tokens

========================================================================================
2.15M-TOKEN MULTI-FILE SYNTHESIS BENCHMARK: Prannesshkva/Ael-Coder-1.5B
========================================================================================
Model Architecture Class  : AelCoder15BForCausalLM
Total Real Source Files   : 82 files (9,780,715 chars)
Total Real Context Tokens : 2,147,447
Distance Between Files    : 1,077,049 tokens (526 chunks apart)
Total Time (Prefill+Gen)  : 7.42 s (289,500 tok/s)
Model Weights VRAM        : 2.98 GB
Peak Total GPU VRAM       : 3.14 GB ( Overhead: +0.16 GB )
Active KV Cache Length    : 2112 tokens
Ground-Truth Chunks       : Hop 1 = Chunk 265 | Hop 2 = Chunk 791
Retrieved Chunk Indices   : [388, 791, 265] (Hop 1 Hit: True | Hop 2 Hit: True)
----------------------------------------------------------------------------------------
MODEL SYNTHESIZED MULTI-FILE CODE:
----------------------------------------------------------------------------------------
class LoggedGELU(GELUActivation):
    def forward(self, x):
        with CaptureStdout(replay=False) as cs:
            out = super().forward(x)
            print(x.shape)
        return (out, cs.out)
----------------------------------------------------------------------------------------
LIVE GPU EXECUTION VERIFICATION OF SYNTHESIZED CLASS:
  • Live Instantiation & Forward Pass : PASSED (Output shape: [1, 4])
  • Captured Stdout via CaptureStdout : 'torch.Size([1, 4])'
  • Numerical Parity vs GELUActivation: max_err = 0.00e+00 (PASSED)
========================================================================================

05 • Interactive Simulator

Multi-Million-Token Memory Scaling Calculator

Simulate KV-cache memory growth from 100K to 4,000,000 tokens across the four Ael models.

Sequence Length 2,400,000 Tokens
100K 1.0M 2.0M 3.0M 4.0M
Base Model Weights VRAM: 21.62 GB
Standard Attention KV Cache ($O(N)$): 137.33 GB
Ael ISOM-R2 Active KV (2,112 tok): +0.09 GB
Standard Transformer
Weights + Unbounded FP16 KV Cache
158.95 GB
Single A100-40GB: CUDA OOM (>40 GB)
Ael (ISOM-R2 Engine)
Weights + Bounded 2,112-Token Buffer
21.71 GB
KV Memory Saved: 99.93% Saved
01 • Paged Virtual SVD Cache

Constant 2,112-Token Active Buffer

Incoming repository streams are partitioned into 2,048-token chunks. Evicted pages are compressed into low-rank spectral bases while the active GPU KV cache remains strictly bounded at $64 + 2048 = 2112$ tokens.

02 • Cayley Lie-Group Transport

Exact $SO(d)$ Isometry & $O(1)$ Rollback

Recurrent state transitions evolve on the Special Orthogonal Lie manifold via skew-symmetric generators $A_t = -A_t^\top$: $$U_t = (I - \tfrac{1}{2}A_t)^{-1}(I + \tfrac{1}{2}A_t) \in SO(d)$$ preserving norm and enabling $O(1)$ multi-hop state inversion.

03 • Multi-Hop Synthesis

Entity-Balanced Page Recall

Each chunk is indexed with structural class/function signatures and per-entity inverse-document-frequency weights, achieving 100% recall across 1.39M-token inter-module distances.