Blockway

OPEN MODEL · PREVIEW · OUT NOW

Agens Volundr

Long context, without the memory bill. Only 18 of its 72 layers keep a KV cache. A dense 32B open model built for tool use, coding and long working sessions, on machines a company can actually own.

  • 32B · dense
  • 72 layers
  • 18 keep a KV cache
  • 262K context
  • Apache-2.0
  • BF16 · INT4
sparse attention · KV cache (17)dense attention · KV cache (1)linear attention · fixed-size state (54)

THE PROBLEM

Long context is a memory problem before it's a compute problem

In a conventional model, every attention layer keeps a KV cache: a copy of what it has seen, which grows with every token. Stack seventy-two of them and, at long context, the cache, not the weights, decides what fits on your machines.

Volundr changes the shape of that curve. 54 layers hold a fixed-size state. Only 18 keep a cache.

context length →cache memory →every layer cachesVolundr: 18 of 72 cache
Schematic, not to scale.

FIVE IDEAS, ONE MODEL

What's inside, and why it matters

KDA · 54 LAYERS

A memory that doesn't grow

Kimi Delta Attention is linear attention: instead of keeping every past token, each layer folds the conversation into a recurrent state of fixed size.

Why it matters: Three quarters of the model carry the whole session forward in the same amount of memory on turn 5 as on turn 500. That's what makes long agent sessions practical on hardware you can buy.

  • fixed-size state
  • no KV cache

TOKENS IN

STATEfixed size

BCSA · 17 LAYERS

Sharp up close, wide at a distance

Blockway Compressed-Sparse Attention reads the latest 4,096 tokens exactly. Everything older is pooled 4:1 into blocks, and a learned indexer picks the 512 most relevant blocks to read.

Why it matters: Every token reads a fixed budget: the recent past in full detail, plus the parts of the distant past that matter. The model can look back across the whole window without paying to look at all of it.

  • 4,096-token exact window
  • far field pooled 4:1
  • top 512 blocks

FAR FIELD · POOLED 4:1 INTO BLOCKS

indexer picks the top 512 blocks

EXACT WINDOW

last 4,096 tokens, every one

DENSE · 1 LAYER

One full view

A single layer of classic full attention closes the stack: here, every token can see every other token.

Why it matters: Fifty-four layers stay lean and seventeen stay selective, and one keeps the complete picture.

  • full attention
  • layer 72

ENGRAM · ADD-ON IN 2 OF THE 72 LAYERS

Memory off the GPU

Engram is a hashed n-gram memory, a large lookup of patterns, and it lives in host RAM rather than on the accelerator.

Why it matters: GPU memory is the scarcest resource in a self-hosted setup. Engram's capacity is paid for in ordinary system RAM, which is far cheaper and far larger.

  • hashed n-gram memory
  • held in host RAM
  • attached at 2 of the 72 layers

GPU MEMORY

model weights and caches

HOST RAM

Engram n-gram memory

mHC · 4 STREAMS

Four lanes instead of one

A conventional transformer passes information between layers along one residual stream. Volundr uses four in parallel.

Why it matters: More room to carry information through 72 layers of depth, so the model doesn't have to squeeze everything into a single path.

  • four residual streams

DENSE, NOT MIXTURE-OF-EXPERTS

What you see is what you run

About 32B parameters, and every layer runs on every token. No hidden expert pool to keep resident, so the footprint is exactly what it says on the label, and capacity planning for your own machines stays simple.

AVAILABLE NOW

What ships

BF16

full precision

Runs on

two 48 GB GPUs, or one H200

Blockway/Agens-Volundr-32B-Preview

INT4

31.7 GiB

Runs on

one 48 GB GPU

Blockway/Agens-Volundr-32B-Preview-INT4

262K-token context

54 of 72 layers hold a fixed-size state, so long sessions stay within reach

Apache-2.0 open weights

on Hugging Face, for commercial use

Serving

Blockway's sglang build (github.com/BlockWayz/agens-sglang); stock sglang and vLLM can't load it yet

GitHub

github.com/BlockWayz/Agens-Volundr: model card, benchmarks and links

GGUF / llama.cpp

planned

Model card, evaluation settings and known limitations are on Hugging Face.

PERFORMANCE

Long context, same speed

Only 18 of 72 layers keep a KV cache, so there is little to re-read as the conversation grows. From 8K to 128K tokens of context, decode drops by 0.2 tok/s.

Decode, tok/s

  • 1K25.1
  • 8K24.1
  • 32K24.1
  • 64K24.0
  • 128K23.9

BF16 on two 48 GB GPUs, single user, Blockway sglang build. INT4 on one 48 GB GPU: 29.1 tok/s at 32K.

Benchmarks

Every model run by Blockway on one harness, with the same settings. Volundr is ahead of Qwen3.8-27B on LiveCodeBench v6, HumanEval, AIME 2025 and MATH-500. Agent tasks trail in this Preview; closing that gap is the focus of the full v1.

Benchmark table: Agens Volundr 32B Preview, Qwen3.8-27B and Agens Pilot on 13 benchmarks, with footnote

Models we train. Hardware you own.

Agens Volundr 32B Preview is out. Download the weights on Hugging Face, and tell us where it breaks.