hanzo-engine

Hanzo Engine is Hanzo AI's high-performance, Rust-native inference and embedding engine that powers the entire Hanzo ecosystem.

Hanzo Engine - Native Rust Inference & Embedding Engine

Category: Hanzo Ecosystem Skill Level: Intermediate to Advanced Prerequisites: Rust knowledge helpful but not required for API usage Related Skills: hanzo-node.md, python-sdk.md, go-sdk.md, zenlm.md

Overview

Hanzo Engine is Hanzo AI's high-performance, Rust-native inference and embedding engine that powers the entire Hanzo ecosystem. Built on mistral.rs with Hanzo-specific optimizations, it provides blazing-fast local inference for all ZenLM models and industry-standard LLMs.

Core Philosophy: Maximum performance, native Rust implementation, multimodal support, and seamless integration with Hanzo Node and Cloud.

Key Features

๐Ÿš€ Blazingly Fast Performance

๐Ÿ”ฎ All-in-One Multimodal

๐ŸŽฏ Embeddings First-Class

๐ŸŒ Multiple APIs

๐Ÿ”— Embedded Everywhere

Architecture

Hanzo Engine in the Stack

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ ZenLM Models (zen-nano, zen-eco) โ”‚
โ”‚ zen-agent, zen-musician, zen-thinking, etc. โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
 โ†“
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ Hanzo Engine (Rust - port 36900) โ”‚
โ”‚ โ”œโ”€ Native inference & embeddings โ”‚
โ”‚ โ”œโ”€ Optimized for Qwen3 models (#1 MTEB) โ”‚
โ”‚ โ”œโ”€ Multimodal: text, vision, audio โ”‚
โ”‚ โ”œโ”€ PagedAttention, FlashAttention, ISQ, MLX โ”‚
โ”‚ โ””โ”€ MCP support built-in โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
 โ†“
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ Embedded in: โ”‚
โ”‚ โ”œโ”€ Hanzo Node (local inference) โ”‚
โ”‚ โ””โ”€ Cloud Nodes (distributed inference) โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
 โ†“
โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ SDKs connect to Node/Cloud: โ”‚
โ”‚ โ”œโ”€ Python SDK (hanzoai package) โ”‚
โ”‚ โ”œโ”€ Go SDK (github.com/hanzoai/go-sdk) โ”‚
โ”‚ โ””โ”€ JavaScript/TypeScript SDK (@hanzo/sdk) โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Based on mistral.rs

Hanzo Engine extends mistral.rs with:

Upstream Sync: Fully synchronized with mistral.rs (commit 530463af1)

Installation

From Source (Recommended)

# Clone Hanzo Engine
git clone https://github.com/hanzoai/engine.git
cd engine

# Build for macOS (Metal backend)
cargo build --package hanzo-engine --release --no-default-features --features metal

# Build for Linux (CUDA backend)
cargo build --package hanzo-engine --release --features cuda

# Install binary
cargo install --path hanzo-engine --no-default-features --features metal

Via Cargo

# Install from GitHub
cargo install --git https://github.com/hanzoai/engine hanzo-engine

# With CUDA support (Linux)
cargo install --git https://github.com/hanzoai/engine hanzo-engine --features cuda

Verify Installation

hanzo-engine --version
hanzo-engine --help

Quick Start

1. Pull a Model

# Pull ZenLM model (zen-eco-4B)
hanzo-engine pull qwen/qwen3-4b

# Pull embedding model (Qwen3-Embedding-8B)
hanzo-engine pull qwen/qwen3-embedding-8b

# Pull from Ollama
hanzo-engine pull ollama://llama3.2:3b

# Pull GGUF model from URL
hanzo-engine pull https://huggingface.co/TheBloke/Llama-2-7B-GGUF/resolve/main/llama-2-7b.Q4_K_M.gguf

2. Start the Server

# Start on default port (36900)
hanzo-engine serve

# With specific port
hanzo-engine serve --port 36900

# With specific model
hanzo-engine serve --model qwen3-4b

# With custom model directory
hanzo-engine serve --model-dir ~/.hanzo/models

3. Test Inference

# Chat completions (OpenAI-compatible)
curl -X POST http://localhost:36900/v1/chat/completions \
 -H "Content-Type: application/json" \
 -d '{
 "model": "qwen3-4b",
 "messages": [{"role": "user", "content": "Explain Rust ownership"}],
 "temperature": 0.7
 }'

# Embeddings
curl -X POST http://localhost:36900/v1/embeddings \
 -H "Content-Type: application/json" \
 -d '{
 "model": "qwen3-embedding-8b",
 "input": "Hello, Hanzo Engine!"
 }'

ZenLM Native Support

All ZenLM models are natively supported in Hanzo Engine:

zen-nano (0.6B)

hanzo-engine pull qwen/qwen3-0.6b
hanzo-engine serve --model qwen3-0.6b

zen-eco (4B) - instruct/thinking/agent variants

# zen-eco-instruct
hanzo-engine pull qwen/qwen3-4b

# zen-eco-thinking (with chain-of-thought)
hanzo-engine pull qwen/qwen3-4b-cot

# zen-agent (tool calling & function execution)
hanzo-engine pull qwen/qwen3-4b-function-calling

zen-musician (7B) - Music generation with lyrics

hanzo-engine pull map/yue-s1-7b-anneal-en-cot

zen-vision (Vision-language)

# Qwen3-VL for vision understanding
hanzo-engine pull qwen/qwen3-vl-8b

Performance: All ZenLM models run at 50+ tokens/sec on consumer hardware (M1, RTX 3060).

Integration with Hanzo Node

Hanzo Engine is the core inference backend for Hanzo Node:

Configuration

# In Hanzo Node configuration (.env or environment)
export EMBEDDINGS_SERVER_URL="http://localhost:36900" # Engine port
export EMBEDDING_MODEL_TYPE="qwen3-embedding-8b"
export USE_NATIVE_EMBEDDINGS="true"
export HANZO_ENGINE_ENABLED="true"

Automatic Connection

When you start Hanzo Node, it automatically connects to Hanzo Engine:

# Start Hanzo Engine
hanzo-engine serve --port 36900

# Start Hanzo Node (connects to Engine on port 36900)
hanzo-node start --mine --gpu

Hanzo Node will:

Python SDK Integration

from hanzo import Hanzo

# Connect to local Hanzo Engine (via Hanzo Node or directly)
hanzo = Hanzo(
 inference_mode='local',
 node_url='http://localhost:8080' # Hanzo Node (which uses Engine)
)

# Or connect directly to Engine
hanzo_engine = Hanzo(
 base_url='http://localhost:36900',
 api_key='not-needed-for-local'
)

# Chat completion (uses zen-eco via Engine)
response = hanzo_engine.chat.completions.create(
 model='qwen3-4b',
 messages=[{'role': 'user', 'content': 'Explain Rust'}]
)

# Embeddings (uses Qwen3-Embedding via Engine)
embedding = hanzo_engine.embeddings.create(
 model='qwen3-embedding-8b',
 input='Hello, Hanzo Engine!'
)

print(f"Embedding dimensions: {len(embedding.data[0].embedding)}") # 4096

Go SDK Integration

package main

import (
 "context"
 "fmt"
 "github.com/hanzoai/go-sdk"
 "github.com/hanzoai/go-sdk/option"
)

func main() {
 // Connect to Hanzo Engine via Hanzo Node
 client := hanzoai.NewClient(
 option.WithInferenceMode("local"),
 option.WithNodeURL("http://localhost:8080"),
 )

 // Chat completion
 response, err := client.Chat.Completions.Create(
 context.Background(),
 hanzoai.ChatCompletionCreateParams{
 Model: hanzoai.F("qwen3-4b"),
 Messages: hanzoai.F([]hanzoai.ChatCompletionMessageParam{
 hanzoai.UserMessage("Explain Rust ownership"),
 }),
 },
 )

 if err != nil {
 panic(err)
 }

 fmt.Println(response.Choices[0].Message.Content)
}

Rust Native API

For maximum performance, use the native Rust API:

use mistralrs::{
 ChatCompletionRequest, IsqType, Loader, MistralRs, ModelDType,
 NormalLoaderBuilder, NormalRequest, Request, RequestMessage, Response,
 SamplingParams, SchedulerConfig, TokenSource
};
use tokio::sync::mpsc::channel;

#[tokio::main]
async fn main() -> anyhow::Result<()> {
 // Load ZenLM model (zen-eco-4B)
 let loader = NormalLoaderBuilder::new(
 "qwen/qwen3-4b",
 None,
 None,
 Some(ModelDType::Auto),
 )
 .build();

 let model = loader.load_model().await?;
 let pipeline = model.build_pipeline(SchedulerConfig::default_config())?;

 let mistralrs = MistralRs::new(pipeline)?;

 // Create request
 let (tx, mut rx) = channel(10_000);
 let request = Request::Normal(NormalRequest {
 messages: RequestMessage::Chat(vec![
 ChatCompletionRequest {
 role: "user".to_string(),
 content: "Explain Rust ownership".to_string(),
 }
 ]),
 sampling_params: SamplingParams::default(),
 response: tx,
 ..Default::default()
 });

 // Send request
 mistralrs.get_sender().send(request).await?;

 // Receive response
 let response = rx.recv().await.unwrap();
 println!("{}", response.choices[0].text);

 Ok(())
}

Model Management

Pull Models

# From HuggingFace
hanzo-engine pull qwen/qwen3-4b

# From Ollama
hanzo-engine pull ollama://llama3.2:3b

# From MLX Community
hanzo-engine pull mlx://mlx-community/Llama-3.2-3B-Instruct-4bit

# From direct URL
hanzo-engine pull https://example.com/model.gguf

List Downloaded Models

hanzo-engine list

# Output:
# Downloaded models:
# - qwen3-4b (4.3 GB) - /Users/z/.hanzo/models/qwen3-4b
# - qwen3-embedding-8b (8.1 GB) - /Users/z/.hanzo/models/qwen3-embedding-8b
# - llama3.2-3b (3.2 GB) - /Users/z/.hanzo/models/llama3.2-3b

Model Storage

Default model directory: ~/.hanzo/models/

Custom directory:

export HANZO_MODELS_DIR=/path/to/models
hanzo-engine serve --model-dir /path/to/models

Performance Optimizations

PagedAttention

Memory-efficient attention with dynamic memory allocation:

# Enable PagedAttention (default for long contexts)
hanzo-engine serve --paged-attention

# Adjust memory usage
hanzo-engine serve --gpu-memory-fraction 0.9

Benefits:

FlashAttention

Ultra-fast attention computation:

# FlashAttention is automatic with CUDA
cargo build --release --features "cuda flash-attn"

# Verify FlashAttention is active
hanzo-engine serve --log-level debug
# Look for: "Using FlashAttention for faster inference"

Benefits:

In-Situ Quantization (ISQ)

On-the-fly model quantization for lower memory:

# Quantize to 4-bit on load
hanzo-engine serve --isq Q4K

# Quantize to 8-bit
hanzo-engine serve --isq Q8_0

# Available formats: Q4K, Q4_0, Q5K, Q8_0, Q8_1

Benefits:

MLX (Apple Silicon)

Optimized for M1/M2/M3 Macs:

# Build with Metal backend (macOS)
cargo build --package hanzo-engine --release --no-default-features --features metal

# Use MLX models for maximum performance
hanzo-engine pull mlx://mlx-community/Llama-3.2-3B-Instruct-4bit
hanzo-engine serve --model llama-3.2-3b-instruct-4bit

Benefits on M1 Max:

MCP Support

Hanzo Engine has built-in MCP support for agentic workflows:

# Start Engine with MCP enabled
hanzo-engine serve --mcp-enabled --mcp-port 3691

# MCP tools automatically exposed:
# - hanzo_infer: Run inference on any model
# - hanzo_embed: Generate embeddings
# - hanzo_list_models: List available models

MCP Integration Example

import { MCPClient } from '@hanzo/mcp'

const client = new MCPClient({
 host: 'localhost',
 port: 3691
})

// Use Hanzo Engine tools via MCP
const result = await client.callTool('hanzo_infer', {
 model: 'qwen3-4b',
 prompt: 'Explain MCP',
 temperature: 0.7
})

console.log(result.response)

Supported Models

Embedding Models (Optimized)

Reranker Models

LLM Models

All models from mistral.rs:

Vision-Language Models

Hardware Requirements

Minimum (zen-nano 0.6B)

Recommended (zen-eco 4B)

Optimal (zen-musician 7B)

Performance Benchmarks

| Model | Hardware | Tokens/sec | Latency | |-------|----------|------------|---------| | zen-nano 0.6B | M1 8GB | 100+ | 10ms | | zen-eco 4B | M1 16GB | 50+ | 20ms | | zen-eco 4B | RTX 3060 12GB | 45+ | 22ms | | zen-musician 7B | M1 Max 32GB | 35+ | 28ms | | zen-musician 7B | RTX 3090 24GB | 60+ | 16ms |

Troubleshooting

Engine won't start

# Check port availability
lsof -i :36900

# Use different port
hanzo-engine serve --port 37000

# Check logs
hanzo-engine serve --log-level debug

Model not found

# Verify model is downloaded
hanzo-engine list

# Re-pull model
hanzo-engine pull qwen/qwen3-4b --force

# Check model directory
ls ~/.hanzo/models/

Out of memory

# Use quantized model
hanzo-engine serve --model qwen3-4b --isq Q4K

# Reduce GPU memory usage
hanzo-engine serve --gpu-memory-fraction 0.8

# Use smaller model
hanzo-engine pull qwen/qwen3-0.6b

Slow inference

# Enable FlashAttention (CUDA)
cargo build --release --features "cuda flash-attn"

# Enable PagedAttention
hanzo-engine serve --paged-attention

# Use quantized model for faster loading
hanzo-engine pull qwen/qwen3-4b-q4k

Hanzo Engine vs Alternatives

| Feature | Hanzo Engine | llama.cpp | vLLM | Ollama | |---------|--------------|-----------|------|--------| | Language | Rust | C++ | Python | Go | | Performance | โญโญโญโญโญ | โญโญโญโญ | โญโญโญโญโญ | โญโญโญโญ | | Multimodal | โœ… All | โœ… Limited | โœ… Yes | โœ… Yes | | Embeddings | โœ… Optimized | โŒ No | โœ… Yes | โœ… Yes | | MCP Support | โœ… Native | โŒ No | โŒ No | โŒ No | | PagedAttention | โœ… Yes | โŒ No | โœ… Yes | โŒ No | | FlashAttention | โœ… Yes | โœ… Limited | โœ… Yes | โŒ No | | Apple Silicon | โœ… MLX | โœ… Metal | โŒ No | โœ… Metal | | Model Management | โœ… Built-in | Manual | Manual | โœ… Built-in | | OpenAI API | โœ… Yes | โœ… Yes | โœ… Yes | โœ… Yes | | Hanzo Integration | โœ… Native | โŒ No | โŒ No | โŒ No |

Why Hanzo Engine?

Related Skills

Additional Resources


Remember: Hanzo Engine is the native Rust inference and embedding engine powering the entire Hanzo ecosystem. All ZenLM models run natively with optimal performance through Engine's advanced features like PagedAttention, FlashAttention, and ISQ.