EuroHPC Grant — Training Concept¶
Proposal No. EHPC-DEV-2026D06-XXX 4,500 Node-Hours · 18,000 GPU-Hours · AMD MI250X · ROCm Stack · 2 TB Storage · 6 Months
Official Development Access Call Award notice received on 2026-06-05.
Overview¶
This document outlines how to use the EuroHPC compute grant to replace or extend deterministic routing and internal process management in the MoE Sovereign Orchestrator. The trained model inference must run primarily on CPUs (legacy hardware) with maximal efficiency, with optional GPU acceleration.
1. Highest Leverage Point: Where AI Beats Deterministic Logic¶
Ranked by impact, based on current codebase analysis:
| Rank | Component | File | Current Logic | Problem |
|---|---|---|---|---|
| 1 | planner_node |
graph/planner.py:153 |
Full LLM call → JSON plan | 2–5s latency, remote LLM dependency, hits 100% of non-cache requests |
| 2 | estimate_complexity |
complexity_estimator.py:127 |
5 regex patterns + zlib AIC | No semantic disambiguation; "Solve x=2" → complex (wrong) |
| 3 | semantic_router_node |
graph/router_nodes.py:227 |
ChromaDB cosine on prototype queries | Hard thresholds ROUTE_THRESHOLD=0.18, no domain-specific fine-tuning |
| 4 | _detect_query_temperature |
graph/planner.py:140 |
2 regex → | Three-step discretization, no continuous calibration |
| 5 | _infer_tier |
services/inference.py:689 |
Regex on model name | No performance-history feedback in tier decision |
Primary lever: planner_node distillation. The LLM planner produces a structured decision
on every request: Query → JSON plan with [{"task": "...", "category": "..."}]. This is a
clearly bounded, learnable mapping. A distilled SLM (1–2B parameters) can handle this call
in <100ms locally on CPU — without network, without GPU.
2. Non-LLM & SLM Alternatives — Code-Based Analysis¶
2a. complexity_estimator.py → DeBERTa-v3-small Encoder Classifier¶
Current logic: _COMPLEX_MARKERS, _MEMORY_RECALL_RE, _RESEARCH_MARKERS (regex)
+ _aic_compressibility (zlib, Kolmogorov proxy)
+ thresholds: _TRIVIAL_TOKEN_MAX=15, _COMPLEX_TOKEN_MIN=80
Problem: No semantic disambiguation.
"Solve the problem" → complex (keyword hit), but may be trivial.
"What is the best strategy for X given Y and Z" → moderate (no keyword hit)
Replacement: DeBERTa-v3-small (22M parameters) as a 4-class classifier
{trivial, moderate, complex, memory_recall}.
- Inference: ONNX INT8 → 3–6ms on x86_64 CPU
- Training: 50K annotated queries (frontier LLM generated + GAIA logs) → 4 GPU-hours
- Advantage over regex: understands that "solve x=2" ≠ "solve partial differential equation"
2b. semantic_router_node + Prototype ChromaDB → Fine-tuned Bi-Encoder¶
Current logic: router_nodes.py:98-107 — upsert _ROUTE_PROTOTYPES into ChromaDB
router_nodes.py:254 — ROUTE_THRESHOLD=0.18, ROUTE_GAP=0.10 hard-coded
Problem: General-purpose embeddings (ChromaDB default) for domain-specific routing.
No feedback loop: thresholds never adapt.
Replacement: paraphrase-multilingual-MiniLM-L12-v2 (22M parameters) fine-tuned with
Triplet Loss on (query, positive_category, negative_category) pairs.
- ONNX FP16 → 8–12ms embedding + FAISS KNN
- Replaces ChromaDB call and hard thresholds with a learned decision space
- No threshold tuning required — classifier outputs classes directly
- Multilingual stability: German queries match English prototypes reliably
2c. fuzzy_router_node + routing_bandit.py → Lightweight RL Policy (Decision MLP)¶
Current logic:
fuzzy_router_node: _compute_routing_confidence → Gödel T-norm → threshold gates
routing_bandit.py: Beta-Bernoulli Thompson Sampling per (gate, context) bucket
Cold-start problem: ROUTING_BANDIT_MIN_DATAPOINTS = 10 → falls back to heuristic
Problem: 4 Redis calls (2 arms × 2 gates) per request, no global context
Replacement: 3-layer MLP (256→128→64→4) as offline RL policy.
- Input state:
[complexity_onehot(4), query_embed(32), cache_distance(1), node_load(N), hour_of_day(1)] - Output: 4-dim discrete
{(research=on/off) × (graphrag=on/off)} - Reward:
judge_confidence_score × (1 - latency_normalized) - cost_per_expert - ONNX FP32 → <1ms on CPU (vs. 4+ async Redis calls)
- Training: offline RL on existing benchmark logs from Grafana/Prometheus
2d. planner_node → Distilled SLM (primary recommendation)¶
Current logic: planner.py:490-558 — 500+ lines prompt construction → LLM → JSON parse
PLANNER_RETRIES=2, Valkey cache for repeated queries
Fast-paths: trivial (direct), memory_recall (direct), agent mode (direct)
Remaining: moderate + complex → always LLM call
Replacement: phi4:14b fine-tuned as structured JSON planner.
- Model selection rationale: phi4:14b is the only sub-35B model that achieves
planner_ok=truein the MoE Sovereign benchmark without any fine-tuning (latency 36.1s, benchmark 2026-07). After SFT on 200K high-quality planner pairs it is expected to significantly outperform granite4.1:3b post-SFT in routing precision and task-description quality. granite4.1:3b remains as ultra-lightweight ablation and fallback for severely resource-constrained deployments (≤4 GB RAM). - Learned task:
(Query, ExpertCategories, ToolDesc) → [{task, category, mcp_tool?, search_query?}] - GGUF Q4_K_M → ~10 tok/s on x86_64 with AVX2 (llama.cpp); ~25 tok/s with AVX-512 BF16
- At ~50–100 planner output tokens: 500–1000ms CPU inference vs. 2–5s remote GPU
2e. _infer_tier + _select_node → XGBoost Ranker¶
Replacement: XGBoost Ranker on (model_name_embed, category, complexity, node_load,
vram_free, requests_pending) → expected_confidence.
- ONNX export → 0.3ms, eliminates several async calls per request
3. Training Strategy: 18,000 GPU-Hours Optimally Used¶
ROCm stack on MI250X: vLLM ≥0.6.0 with ROCm 6.x, PyTorch 2.4+, Flash-Attention-2 (Composable Kernel for ROCm). MI250X has 2×64 GB HBM2e (128 GB total) per node → trains 70B models in float16 without offloading.
Phase 1 — Synthetic Data Generation (Months 1–2, 2,500 GPU-h)¶
| Dataset | Method | Volume | GPU-Hours |
|---|---|---|---|
| Query → Complexity labels | Llama-3.3-70B batch inference on MI250X | 500K pairs | 600 h |
| Query → Expert-category labels | Llama-3.3-70B + existing template definitions | 300K pairs | 400 h |
| (Query, Plan-JSON, Judge-Score) triples | Llama-3.1-405B as primary teacher (Q4_K_M on 1 node, 10/10 quality), Qwen3-235B-A22B as fallback (9/10, faster) | 200K triples | 350 h |
| (State, Action, Reward) for RL | Replay benchmark logs with simulated judge reward | 500K transitions | 200 h |
| DPO preference pairs (planner) | Two plans per query → judge ranks, builds preferred/rejected | 100K pairs | 500 h |
Quality assurance: Each data batch validated by a separate frontier judge (judge_prompt
pattern from templates). GAIA-2023-Validation held out as untouched test set.
Phase 2 — Encoder/Classifier Training (Month 2, 800 GPU-h)¶
- Complexity classifier: DeBERTa-v3-small, HuggingFace Trainer, bfloat16, lr=3e-5, 10 epochs → 200 GPU-h
- Routing classifier (semantic router):
paraphrase-multilingual-MiniLM-L12-v2, Triplet Loss, batch=512 → 300 GPU-h - Reward model (for Phase 5): DeBERTa-v3-large fine-tuned on (Plan, GAIA-Correct: bool) → 300 GPU-h
Phase 3 — Planner SFT + DPO (Months 2–4, 8,000 GPU-h)¶
Base model: phi4:14b (only sub-35B model with proven planner_ok=true without fine-tuning;
benchmark 2026-07: planner_ok=true, latency=36.1s base / expected <1s post-SFT)
Context: 4096 tokens (planner prompt ~2000–3000 tokens + output ~200 tokens)
Batch size: 64 (bfloat16, gradient checkpointing, ZeRO-3)
SFT Phase 1 (teacher forcing on planner outputs):
Dataset: 200K (Prompt, JSON-Plan) pairs from Llama-3.1-405B teacher
(Q4_K_M on 1 LUMI-G node, ~300 GPU-h accounted in Phase 1)
Hard-reject filter: arithmetic→precision_tools, §-query pattern,
research-before-code — all violations auto-rejected before training
3 epochs, lr=1e-4, cosine schedule with warmup, AdamW
Budget: 4,500 GPU-h
DPO Phase 2 (Direct Preference Optimization):
Dataset: 100K (Prompt, Preferred-Plan, Rejected-Plan) triples
Rejected = plans that failed hard_reject filter (Phase 1 generation)
β=0.1, lr=5e-5, 2 epochs
Budget: 2,200 GPU-h
Ablations (model sizes, data quality):
granite4.1:3b as ultra-lightweight ablation (CPU ≤4 GB RAM target) → 800 GPU-h
granite4.1:8b as intermediate ablation → 500 GPU-h
Target metric: GAIA plan quality score ≥95% of Llama-3.1-405B teacher at <1/50 inference cost. phi4:14b CPU fallback target: complete plan in <1s on Xeon Skylake (AVX-512 BF16).
Ist-Abgleich (2026-07-24, LUMI-G Job 20190726): Basismodell abweichend von diesem Konzept auf
Qwen3-8Bfestgelegt (bewusste Entscheidung, nichtphi4:14b— bessere ROCm/DeepSpeed-ZeRO-3-Unterstützung, sieheeurohpc_lumi_activity_report.mdAktivität 10). Kritischer Konfigurationsfehler: Der tatsächliche Trainingslauf setztemax_seq_len=1536statt der hier korrekt dokumentierten 4096 Token — der Systemprompt allein benötigt bereits ~2.500 Token, wodurch die Ziel-JSON-Antwort bei jedem der 259.829 Trainingsbeispiele außerhalb des Loss-Fensters lag (100% Truncation-Rate laut Trainingslog). Der Fine-Tune war dadurch strukturell wirkungslos, obwohl der Job formal erfolgreich durchlief. Root-Cause-Analyse und Fix:eurohpc_lumi_activity_report.mdAktivität 10. Merke für zukünftige Phasen (4/5): Die in diesem Dokument vorab kalkulierten Context-Werte sind bindend — Abweichungen im tatsächlichen Trainingsscript müssen vor Job-Submission gegen dieses Konzept geprüft werden, nicht erst nachträglich anhand der Logs.
Phase 4 — Offline RL Routing Policy (Months 4–5, 3,500 GPU-h)¶
Method: Decision Transformer (offline RL) on historical benchmark logs
Alternative: Behavior Cloning + Advantage Weighted Regression (AWR)
State: [complexity_embed(32), query_embed(32), cache_distance(1), load_vector(N), bandit_context(8)]
Action: 4-dim discrete {research, graphrag} × {on, off}
Reward: judge_confidence_score × (1 - latency_normalized) - 0.1 × cost_per_expert_call
Training: 3 epochs on 500K (State, Action, Reward, Next-State) transitions
Budget: 2,000 GPU-h (Transformer) + 1,500 GPU-h (ablations, online eval on GAIA)
Phase 5 — Planner RLHF Fine-Tuning (Months 5–6, 2,200 GPU-h)¶
Reward model from Phase 2 (DeBERTa-v3-large) as reward signal
PPO on phi4:14b (best model from Phase 3)
KL-penalty: β=0.05 (conservative, prevents reward hacking)
1 epoch, 50K queries from GAIA + own benchmark queries
Budget: 2,200 GPU-h
Budget Summary¶
| Phase | Content | GPU-Hours |
|---|---|---|
| 1 | Data synthesis (Llama-3.1-405B + other frontier batch inference) | 2,350 |
| 2 | Encoder/classifier + reward model | 800 |
| 3 | Planner SFT phi4:14b + DPO + ablations | 8,000 |
| 4 | Offline RL routing policy | 3,500 |
| 5 | Planner RLHF (PPO on phi4:14b) | 2,200 |
| (consumed) | Judge-v3 training (completed 2026-07) | 437 |
| Total | 17,287 h of 18,000 h ✓ | |
| Puffer | ~713 h (4 %) |
4. Export and Inference Format for x86_64 Legacy CPUs¶
Target hardware: x86_64 with AVX2 (Haswell+) or AVX-512 (Skylake-SP+), no dedicated inference accelerator.
| Model | Format | Quantization | Compile Flags | CPU Latency | Size |
|---|---|---|---|---|---|
| Planner phi4:14b | GGUF | Q4_K_M (4.3 bit/weight) | llama.cpp: -DLLAMA_AVX512=ON -DLLAMA_AVX512_BF16=ON |
~25 tok/s (Xeon w/ AVX-512 BF16) | ~8.0 GB |
| Planner phi4:14b (GPU) | GGUF | Q4_K_M | llama.cpp CUDA/ROCm backend | ~200 tok/s | ~8.0 GB |
| Planner granite4.1:3b (fallback) | GGUF | Q4_K_M | same flags | 15–25 tok/s | ~2.0 GB |
| Complexity classifier | ONNX | INT8 Dynamic (22M params) | ORT 1.18+ CPUExecutionProvider, AVX2 | 3–6 ms | ~22 MB |
| Semantic router embedding | ONNX | FP16 | ORT with graph_optimization_level=ORT_ENABLE_ALL |
8–12 ms | ~44 MB |
| RL routing policy MLP | ONNX | FP32 (trivial, <1M params) | Standard ORT | <1 ms | <5 MB |
| XGBoost node ranker | ONNX | FP32 | XGBoost ≥2.0 native ONNX export | <0.3 ms | <2 MB |
Quantization Rationale¶
- Q4_K_M (GGUF): K-Quant with mixed 4-bit/6-bit quantization in critical layers. Retains ~97–98% model quality at 4× compression. Standard for llama.cpp CPU deployment.
- INT8 Dynamic (ONNX): Weights INT8, activations FP32. No calibration dataset required. 2–3× speedup on AVX2 vs FP32 via VNNI instructions on Intel CPUs.
- FP16 for embeddings: INT8 quant of sentence encoders measurably degrades cosine similarity (ΔnDCG ~2–4%); FP16 is the safe compromise.
OpenVINO as Alternative for Intel Xeon Fleets¶
- ONNX → OpenVINO IR via
mo --input_model model.onnx --data_type FP16 - INT8 via POT (Post-Training Optimization Toolkit) without quality loss on encoder models
- Latency advantage: +30–50% over ORT on Intel Xeon Scalable (Cooper Lake with AVX-512 VNNI)
5. New Capabilities After Successful Training¶
| # | Capability | Current Code Limitation | New Capability |
|---|---|---|---|
| 1 | Offline planner without remote LLM | planner.py:566 — _invoke_llm_with_fallback to external LLM |
GGUF planner runs locally on CPU, no network, no GPU node. Enables air-gap deployment |
| 2 | Planner latency <100ms | 2–5s on remote call; 1800s Valkey cache TTL as workaround | Local GGUF inference: 200–500ms for complete plan. Cache dependency drops drastically |
| 3 | Semantic complexity recognition | complexity_estimator.py — token count + regex only |
Encoder distinguishes "solve x=2" (trivial) from "solve nonlinear PDE" (complex), keyword-independent |
| 4 | Cold-start-free adaptive routing | routing_bandit.py:87 — ROUTING_BANDIT_MIN_DATAPOINTS guard |
RL policy generalizes to new query types via transfer, no minimum data volume before activation |
| 5 | Continuous policy updates without redeployment | Thompson bandit is stateless except Redis counters; new categories = new cold starts | RL policy can be incrementally fine-tuned with new (State, Reward) data (LoRA) |
| 6 | Routing cost-quality Pareto optimization | No explicit cost model in routing | RL agent minimizes latency + cost_per_token under quality constraint from judge feedback |
| 7 | Multilingual semantic routing stability | ChromaDB prototypes are language-dependent; German queries match English prototypes poorly | paraphrase-multilingual-MiniLM-L12-v2 fine-tuned → equal routing quality on DE/EN/FR |
| 8 | Confidence score per plan | Planner LLM outputs no confidence signal | Distilled SLM can quantify planning uncertainty via temperature sampling + token log-probs → fallback trigger when conf < threshold |
What Does Not Change (Intentional Non-Substitution)¶
services/routing.py— template binding remains deterministic (admin-configured, no ML)pipeline/logic_types.py— fuzzy/paraconsistent types remain as formal correctness guarantees_select_nodesticky session logic — remains deterministic for UX consistency- Judge node — remains large LLM (35B+); quality bottleneck is not compressible
Configuration Switch Design¶
Following the existing pattern of TRIVIAL_FAST_PATH_ENABLED (planner.py:223), each new
AI component should be gated by a feature flag:
Hybrid mode: SLM as first pass, LLM fallback when confidence < threshold. No hard switch.
6. Dynamic Template Synthesizer¶
Concept¶
Current system: Template = user-profile binding
New (optional, feature-flagged): Template = prompt-adaptive synthesis
The synthesizer does not replace static templates but extends them as an optional mode. Admin configuration overhead for special cases is eliminated.
Code Integration Point¶
Current flow:
main.py → _resolve_user_experts() → user_experts dict → stream_response() → Graph
↑ loaded statically from DB
New flow with synthesizer:
main.py → [Dynamic Template Synthesizer] → user_experts dict → stream_response() → Graph
↑
Prompt-adaptive generation
(replaces or overlays DB template)
Concrete code location (main.py:1283–1347):
# Today:
user_experts = _resolve_user_experts(permissions_json, override_tmpl_id, ...)
# With synthesizer (optional path):
if TEMPLATE_SYNTHESIS_ENABLED and not override_tmpl_id:
user_experts = await _synthesize_template(user_input, available_categories)
else:
user_experts = _resolve_user_experts(permissions_json, ...)
Synthesizer Output Structure¶
Identical structure to _resolve_user_experts() in routing.py:72–116:
{
"legal_advisor": [
{
"model": "qwen3.6:35b",
"endpoint": "n04-rtx",
"forced": true,
"_system_prompt": "Specialist for §613a BGB business transfers and employment law. Focus on dismissal protection and collective agreement continuity.",
"_tier": 2
}
],
"code_reviewer": [
{
"model": "qwen3-coder:30b",
"endpoint": "n04-rtx",
"_system_prompt": "Python FastAPI Security Review. Check for OWASP Top 10, async patterns, dependency injection vulnerabilities.",
"_tier": 2
}
]
}
Additionally synthesized (mapped to _resolve_template_prompts() output, routing.py:128–220):
{
"planner_prompt": "Prioritize §-precise research before implementation. Combine legal review + code security.",
"judge_prompt": "Evaluate by: legal correctness, GDPR compliance, code security.",
"enable_graphrag": true,
"enable_web_research": false,
"force_think": true
}
Three Implementation Tiers¶
Tier 1 — Retrieval-Augmented Template Assembly (RATA)¶
No generative training required, deployable immediately
Prompt
↓
DeBERTa classifier → {legal: 0.91, code: 0.78, technical: 0.12}
↓
Threshold filter (>0.5) → ["legal_advisor", "code_reviewer"]
↓
FAISS search: for each category retrieve top-3 semantically similar proven system prompts
from a curated library (curated + GAIA-labelled)
↓
Top-1 system prompt per category → composite template
- Routing flags (
enable_graphrag,web_research) → derived from category classifier (rule-based: legal → graphrag=True) - Latency: 10–25ms (classifier + FAISS)
- No generative model required
Tier 2 — SLM Template Generator (primary EuroHPC target)¶
Generates system prompts adapted to the concrete prompt context
Prompt
↓
phi4:14b fine-tuned → JSON template (fully structured)
↓
Validation (known categories, model endpoints via config)
↓
Composite template
- Model generates contextually refined system prompts instead of generic category defaults
- Latency: 500–1000ms (GGUF Q4_K_M on CPU) — bypassed by cache when query repeated
- Key value-add: Prompt "Explain Docker for a beginner" →
_system_prompt: "Explain step-by-step, use analogies, avoid jargon"vs. "Optimize our Docker Compose security" →_system_prompt: "Senior Security Engineer, focus on non-root user, read-only filesystems, secret management"
Tier 3 — Hierarchical Synthesizer (maximum quality)¶
Prompt
↓
Tier 1: DeBERTa categorizes quickly
↓
Tier 2: SLM refinement of system prompts for selected categories only
↓
Tier 3: Optional RL policy adjusts routing flags based on historical performance
EuroHPC Training Strategy for Synthesizer¶
Training Data Generation (~2,000 GPU-hours re-allocated from Phase 1)¶
Gold standard pairs: (Prompt, Optimal_Template_JSON)
The training pairs must teach the Sovereign-14B model how to generate context-aware, tailored system prompts for the planner, the judge, and every selected expert to optimize orchestration quality.
Prompt: "Our FastAPI app transfers user data to US service providers under Art. 46 GDPR."
→ Optimal template:
{
"planner_model": "qwen-3.6-35b-sovereign@AIHUB",
"judge_model": "llama3.3-70b-ctx4k:latest@N04-RTX",
"planner_prompt": "You are a specialized planner coordinating a legal compliance and code audit. Plan retrieval step first, followed by expert audit tasks.",
"judge_prompt": "You are a GDPR compliance judge. Synthesize legal findings and code recommendations, ensuring strict alignment with Art. 46.",
"enable_graphrag": true,
"enable_web_research": true,
"experts": {
"legal_advisor": {
"system_prompt": "You are a specialist in GDPR Art. 46/47 cross-border transfers. Analyze Transfer Impact Assessments (TIA) and standard contractual clauses.",
"context_window": 32768
},
"code_reviewer": {
"system_prompt": "You are a security code auditor. Audit the FastAPI codebase for encryption in transit, data minimization, and consent validation flows.",
"context_window": 16384
}
}
}
Generation method:
1. Frontier LLM SFT Pair Generation: A frontier LLM (e.g. Llama-3.3-70B on MI250X) will generate 300K (Prompt, Template) training pairs. The generation instructions must enforce creating tailored, prompt-specific system prompts (planner_prompt, judge_prompt, and expert system_prompts) instead of generic defaults.
2. Existing GAIA logs: prompts that led to high judge confidence → extract their effective routing decisions as gold standard
3. Own benchmark logs (Prometheus/Grafana): (input, expert_calls_made, judge_confidence) → reconstruct templates retroactively
DPO preference pairs for system prompt quality:
Prompt: "Explain Kubernetes Networking to a junior developer"
Preferred: _system_prompt = "Explain with analogies (postman = Pod, street = Service),
step-by-step, avoid CIDR/iptables details, use diagrams"
Rejected: _system_prompt = "You are a Kubernetes expert. Answer technical questions."
Evaluation criterion for Preferred/Rejected: judge score after full pipeline execution with the respective preferred vs. rejected system prompt.
Training Budget (within overall Phase 3 budget)¶
| Step | Budget |
|---|---|
| Data synthesis (300K template pairs via frontier LLM) | 800 GPU-h |
| DPO pair generation (100K pairs, judge-scored) | 600 GPU-h |
| SFT phi4:14b on template generation | 2,500 GPU-h |
| DPO fine-tuning | 1,000 GPU-h |
| Ablations (category count, prompt length, retrieval vs. generation) | 500 GPU-h |
| Total | ~4,200 GPU-h |
These 4,200 hours are drawn from the Phase 3 budget — planner SFT and template synthesizer share the granite4.1:3b base checkpoint.
Impact: What Changes Fundamentally¶
| Dimension | Today | With Synthesizer |
|---|---|---|
| System prompt specificity | Generic per category ("You are a code reviewer") | Prompt-contextual ("You are a Python async performance reviewer for FastAPI focusing on DB bottlenecks") |
| Template management overhead | Admin must maintain 50+ templates for every scenario combination | One base template per user group; synthesizer handles fine granularity |
| Cold start for new domains | New domain → admin must build new template | Synthesizer generalizes to unknown domains via transfer learning |
| Multi-domain prompts | Require manually crafted combination templates | Automatically provided with matching, coordinated system prompts per expert |
| User-specific adaptation | Template binding = coarse role segmentation | Synthesizer can incorporate user history from cross-session memory → personalized expert configuration |
| Routing flags (graphrag, web) | Set statically in template | Decided prompt-adaptively by synthesizer |
Relationship to Planner (Synergy, Not Competition)¶
Synthesizer: "WHO" and "HOW" (which experts, with what focus)
Planner: "WHAT" and "IN WHAT ORDER" (which subtasks, which tools)
A distilled planner (Phase 3) and a synthesizer can share the same base checkpoint
phi4:14b — different LoRA adapters for different tasks. This reduces the VRAM
footprint in production deployment: one model (~8 GB Q4), two adapters, sequentially
loaded in <50ms via llama.cpp hot-swap. granite4.1:3b (2 GB) remains available as
ultra-lightweight fallback when RAM < 10 GB.
Appendix: Key Config Constants Referenced¶
| Constant | File | Value | Role in Decision |
|---|---|---|---|
CACHE_HIT_THRESHOLD |
config.py:178 |
0.15 | ChromaDB L1 cache hit gate |
ROUTE_THRESHOLD |
config.py:181 |
0.18 | Semantic router direct-path gate |
ROUTE_GAP |
config.py:182 |
0.10 | Minimum confidence gap for routing |
EXPERT_MIN_SCORE |
config.py:223 |
0.3 | Thompson sampling minimum usable score |
ROUTING_BANDIT_MIN_DATAPOINTS |
config.py |
10 | Cold-start guard for bandit gates |
TRIVIAL_FAST_PATH_ENABLED |
config.py |
False | Reference pattern for new feature flags |
THOMPSON_SAMPLING_ENABLED |
config.py |
True | Expert scoring mode |
JUDGE_REFINE_MAX_ROUNDS |
config.py |
— | Judge refinement loop cap |
PLANNER_RETRIES |
config.py |
2 | Planner JSON parse retry count |
_TRIVIAL_TOKEN_MAX |
complexity_estimator.py:28 |
15 | Word count threshold for trivial |
_COMPLEX_TOKEN_MIN |
complexity_estimator.py:29 |
80 | Word count threshold for complex |
_AIC_TRIVIAL_FLOOR |
complexity_estimator.py:34 |
0.55 | Kolmogorov compressibility → trivial |
_AIC_COMPLEX_CEILING |
complexity_estimator.py:35 |
0.15 | Kolmogorov compressibility → complex |
Appendix B: AMD MI250X vs. NVIDIA CUDA — Practical Field Notes (LUMI-G)¶
This section documents concrete incompatibilities, workarounds, and best practices discovered during active training runs on LUMI-G (AMD MI250X / ROCm). Many standard HuggingFace tutorials implicitly assume NVIDIA/CUDA. These notes close that gap.
Hardware Reference¶
| AMD MI250X (LUMI-G) | NVIDIA A100 SXM | NVIDIA H100 SXM | |
|---|---|---|---|
| Architecture | CDNA2 (GCDs, 2 dies/card) | Ampere | Hopper |
| VRAM / GPU slot | 64 GB HBM2e (per GCD) | 80 GB HBM2e | 80 GB HBM3 |
| VRAM / Full Node | 512 GB (8 GCDs) | 640 GB (8×A100) | 640 GB (8×H100) |
| Peak BF16 TFLOPS | ~383 (per GCD) | ~312 | ~989 |
| Software Stack | ROCm 6.x / HIP | CUDA 12.x | CUDA 12.x |
| EuroHPC Presence | LUMI (Finland, #1 EU) | MareNostrum 5 | Frontier (US) |
Compatibility Matrix (as of 2026-07)¶
| Feature | NVIDIA CUDA | AMD MI250X ROCm | Notes |
|---|---|---|---|
| 4-bit QLoRA (BitsAndBytes NF4) | ✅ Stable | ⚠️ Unstable | ROCm build has sporadic compile errors; avoid for production training. Use BF16 + ZeRO-3 instead. |
| Flash Attention 2 | ✅ Native flash_attn |
❌ Not available | Use attn_implementation="eager". ~10–15% throughput penalty. ROCm FA via Triton exists but unreliable. |
device_map="auto" (Pipeline Parallel) |
✅ Standard | ❌ OOM risk | ROCm layer distribution is unbalanced on heterogeneous model topologies → use device_map=None + ZeRO-3. |
| DeepSpeed ZeRO-3 | ⚠️ Rarely needed | ✅ Required for 35B+ | Primary memory strategy on LUMI-G. Stable. |
| DeepSpeed ZeRO-3 + Multimodal VLMs | ⚠️ Edge case | ❌ Crashes without fix | See "Known Bug" below. Vision-tower zero-param assertion. |
| bfloat16 training | ✅ A100/H100 | ✅ MI250X | Both stable. Preferred over float16 on AMD. |
| RCCL (multi-GPU collective) | ✅ NCCL 2.x | ✅ RCCL (AMD fork) | Functionally equivalent within a node. No measurable difference. |
| mmap of large model files on login nodes | ✅ Works | ❌ Cannot allocate memory |
Login nodes are CPU-only with limited RAM. Load models only in batch jobs with --gpus. |
amdsmi / rocm-smi in containers |
n/a | ⚠️ rsmi_init exception |
Only works when container is bound to GPU resources via SLURM --gpus. Benign warning, training not affected. |
| Transformers new model types | ✅ Immediate support | ⚠️ Container lag 2–4 weeks | qwen3_5_moe required my_venv with a newer Transformers install than the container's /opt/venv. |
Known Bug: DeepSpeed ZeRO-3 + Multimodal VLMs (AMD / All)¶
Symptom: AssertionError: {'status': 'NOT_AVAILABLE', 'numel': 0, ...} on one or more ranks
during the first forward pass.
Root cause: Multimodal models (e.g. Qwen3_5MoeForConditionalGeneration, Qwen2-VL,
InternVL2) contain a vision tower with parameters that are not allocated during text-only
SFT — they have numel=0 and shape=(0,). DeepSpeed ZeRO-3 partitions all registered
parameters, attempts to all_gather these zero-size tensors at runtime, and fails.
Why this surfaces on AMD specifically: On NVIDIA, 4-bit QLoRA (BitsAndBytes) is the standard memory strategy for large models. It does not require ZeRO-3. On AMD/ROCm, BitsAndBytes is unstable, so ZeRO-3 BF16 is used instead — exposing this bug.
Fix (two-part):
# train_judge_lora.py — inside deepspeed.zero.Init() context, BEFORE get_peft_model()
_TEXT_MODULE_PREFIXES = (
"model.language_model", # Qwen3_5MoeForConditionalGeneration
"language_model", # older VL variants
"model.text_model", # Qwen2-VL style
"model.model", # plain CausalLM fallback
"lm_head",
)
for name, param in model.named_parameters():
is_text = any(name.startswith(p) for p in _TEXT_MODULE_PREFIXES)
if not is_text or param.numel() == 0:
param.requires_grad_(False) # prevents ZeRO-3 from partitioning these
// deepspeed_zero3.json
"stage3_param_persistence_threshold": 1e7 // prevents partitioning of small/empty tensors
Recommended AMD MI250X Training Configuration (35B+ models)¶
model = AutoModelForCausalLM.from_pretrained(
base_model_path,
trust_remote_code=True,
attn_implementation="eager", # Flash-Attn 2 unavailable on ROCm
torch_dtype=torch.bfloat16, # BF16 stable; FP16 not recommended on MI250X
device_map=None, # None required for ZeRO-3; "auto" causes OOM
)
// deepspeed_zero3.json — recommended baseline for LUMI-G
{
"bf16": { "enabled": true },
"fp16": { "enabled": false },
"zero_optimization": {
"stage": 3,
"overlap_comm": true,
"contiguous_gradients": true,
"stage3_param_persistence_threshold": 1e7,
"stage3_gather_16bit_weights_on_model_save": true,
"offload_optimizer": { "device": "none" },
"offload_param": { "device": "none" }
}
}
SLURM Script Patterns for LUMI-G¶
# Mandatory for ROCm GPU visibility in Singularity:
export ROCR_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
export HIP_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
export MASTER_ADDR=localhost
export MASTER_PORT=29500
# Use newer Transformers from my_venv if container version is too old:
export SINGULARITYENV_PYTHONPATH=/scratch/.../my_venv/lib/python3.12/site-packages:/opt/venv/lib/python3.12/site-packages
# Launch with torchrun (handles rank/world_size injection):
singularity exec --bind /pfs,/scratch,/projappl,/project "${SIF}" \
torchrun --nproc_per_node=8 --master_addr="${MASTER_ADDR}" --master_port="${MASTER_PORT}" \
train_judge_lora.py --deepspeed deepspeed_zero3.json --no_4bit ...
Note: Do not use
PYTORCH_ALLOC_CONF=expandable_segmentson ROCm — it is not supported and may cause silent memory corruption or training hangs.