Blockway

BLOCKWAY · OPEN MODEL · HONG KONG

Agens Pilot

The Cantonese-first, balanced open model — built on Qwen3.8-27B and post-trained by Blockway to answer in 廣東話, give balanced answers on complex questions, and power our own agent harnesses.

  • 廣東話 first
  • Built on Qwen3.8-27B
  • 1M context
  • Text + Vision
  • Apache-2.0
  • BF16 · FP8 · INT4

What we changed — and what we kept

Qwen3.8-27B is an excellent base. We didn't re-teach it; we gave it the three things a Hong Kong company and its agent products actually needed.

廣東話, first-class

Write in Cantonese, get natural written Cantonese back — Hong Kong usage and register, not a Mandarin answer in disguise. Switches cleanly to 普通話 or English when you do.

Balanced, guardrails kept

Facts and multiple perspectives on complex historical and geopolitical questions — 85% multi-perspective vs 60% for the base — while refusals on universally-harmful requests stay exactly as strict.

Built for our agents

Trained and evaluated against the real workloads of Claway and Codeway: long context, strict instruction-following, reliable tool calls.

Everything Qwen3.8 does

Strong coding (97.0 HumanEval, ~83 LiveCodeBench v6 on our harness), a 1,000,000-token context on a hybrid linear-attention design, and native vision — all inherited from the base.

Reliable tool use

Consistent function calling and multi-step tool workflows — the backbone of dependable agentic behaviour.

Open and honest

Apache-2.0, three official builds, and full attribution to the Qwen team. A focused fine-tune of a base we respect — not a rebrand.

AGENS PILOT VS. ITS BASE

Balanced answers on complex questions

On complex historical and geopolitical questions, Agens presents the facts and multiple perspectives rather than a single viewpoint — and it refuses genuinely harmful requests exactly as strictly as before.

Multi-perspective answers

complex historical & geopolitical questions

Agens Pilot85%
Qwen3.8-27B (base)60%

Single-viewpoint answers

same questions

Agens Pilot10%
Qwen3.8-27B (base)35%

Harmful requests refused

weapons, malware, exploitation — safety retained

Agens Pilot8 / 8
Qwen3.8-27B (base)8 / 8

Same 20 complex-question prompts and 8 harmful prompts to both models, temperature 0.6, single run, 2026-08-29. Balance graded by the base model itself acting as judge — Qwen3.8-27B, thinking off, temperature 0 — on a 0 / 1 / 2 rubric (non-answer / single viewpoint / multi-perspective). Judge script and aggregate scores are published with the model. Small sample; treat as directional.

PURPOSE-BUILT

The engine behind Blockway's agents

Agens Pilot is trained and evaluated against the real workloads of Blockway's own agent stacks — long context, strict instruction-following, reliable tool calls — so it behaves the way an agent should out of the box.

Claway

LIVE

Enterprise AI employee platform

Claway deploys AI employees on a company's own hardware to take on whole roles — accounting, sales, customer service — with private memory of your rules and records. Private deployment, data stays offline, and an AI workforce at roughly 1/5 the cost of human staff. Running in production today.

Visit claway.io

Codeway

NOW OPEN

Coding agent

Repo-scale understanding, edit-and-verify loops and agentic software delivery — a coding harness built to make Agens Pilot's 1M context and code strength count. In your terminal, and as a chat on the web.

Explore Codeway →

General capability

Blockway internal harness — single-sample, thinking enabled, temperature 0.6. The fine-tune changes behaviour, not capability: on general benchmarks Agens tracks its base within measurement noise.

HumanEval (chat)97.0
LiveCodeBench v6~83
GPQA Diamond80.3
RealWorldQA · vision77.2
MathVision · vision65.8
IFBench (strict)61.0

tool-call selection · 16 / 16  ·  figures are internal & conservative — not directly comparable to public leaderboards

VS. RECENT OPEN MODELS · OFFICIAL REPORTED SCORES

GPQA Diamond

Agens Pilot80.3
Llama 4 Scout · Meta57.2
Mistral Small 3.1 · 24B45.4
Gemma 3 · 27B42.4

HumanEval

Agens Pilot97.0
Mistral Small 3.1 · 24B~88

LiveCodeBench

Agens Pilot~83
Llama 4 Scout · Meta32.8
Gemma 3 · 27B29.7

Competitor figures as officially published — Llama 4 Scout, Mistral Small 3.1 & Gemma 3 announcements / technical reports. Cross-harness comparison is approximate: Agens uses Blockway's internal harness with thinking enabled, and LiveCodeBench versions may differ.

Three official builds

Same tokenizer and chat template across all three — they differ only in weight precision.

Quantized builds keep the hybrid linear-attention layers, embeddings and LM head in bf16 for stability — only feed-forward and full-attention weights are quantized. All builds serve in compressed-tensors format via sglang / vLLM. GGUF will follow as soon as upstream llama.cpp supports the Qwen3.5/3.8 hybrid architecture.

Open weights, now on Hugging Face.

Agens Pilot is released under Apache-2.0 in three builds — BF16, FP8 and INT4 — with full attribution to the Qwen team. Grab the weights, or reach out for an enterprise deployment or to build on the Claway and Codeway harnesses.