Zum Inhalt springen

AINode Release Notes

23 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Folge AINode, um die Release Notes in deinen Feed zu holen.

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.33: /v1/decide und /v1/systemone vom modellhaltenden Node

AINode 0.5.33 beantwortet /v1/decide und /v1/systemone nun immer vom Node, der das Modell bereitstellt (inklusive dessen Kalibrierung und Warm-Grammar), und holt bei Bedarf dessen Kalibrierungstabelle per neuem Endpunkt /api/decide/calibration.

The owner answers: a decision request is answered by the node that serves the model, so its calibration and warm grammar apply wherever the request enters, and a slow resolver can no longer freeze a node, which is what stopped the master routing to its peers after the 0.5.32 roll.

Fixed

  • A decision is answered by the node that owns the model, wherever the request enters (#285). /v1/decide and /v1/systemone for a model the receiving node does not serve (the master, usually) are forwarded whole to the serving node's AINode over the fleet key, so the answer carries that node's temperatures and warm grammar instead of calibration: {applied: false, temperatures: null}. Replicas are tried in turn when one is down; a forwarded request is never forwarded again and does not count against the owner's rate limit; with no owner answering, the node calls the engines itself as before. A locally served model is unchanged.
  • When the receiving node does have to call the engines itself, it tempers with the serving node's table (#284). A node with no copy of the model fetches the table from the node it routes to over the fleet key (GET /api/decide/calibration?model=<id>, new) and caches it for five minutes. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.32: Decision-Adapter-Temperaturen und Grammar-Warmup

AINode 0.5.32 wendet bei /v1/decide und /v1/systemone die Temperaturen des Decision-Adapters an, wärmt die Antwortgrammatik beim Binden vor, bewertet Modelle im Decision-Bench auf den öffentlichen Jevals-Sets und erlaubt einen Master ohne Modell.

Decisions use the model's own calibration: /v1/decide and /v1/systemone apply a decision adapter's fitted temperatures and say so in every answer, a decision model warms its answer grammar the moment it binds, the decision bench scores any model on the public Jevals sets with the boards' own formulas, and a master that serves no model is safe to run, which is the first step of moving the fleet's control plane to Atlas.

Added

  • /v1/decide and /v1/systemone apply a decision adapter's own temperatures (#276, #281). When the served model's directory carries temperatures.json, each question's label logprobs are divided by the temperature fitted for its kind (choice, noul, score) before the softmax. Both responses report calibration: {applied, temperatures}, and "calibration": "raw" returns the engine's own spread.
  • A decision model warms its answer grammar as soon as its engine binds (#277, #281). One small constrained question per kind is sent in the background, with each compile time logged. /api/status gains instances[] with warm and warm_compile_seconds per instance. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.20: Neuer Parameter --claude-effort für Harness-Bench

AINode 0.5.20 ergänzt den Parameter ainode-bench harness --claude-effort <level>, mit dem sich der Reasoning-Aufwand von Claude Code im Harness-Bench festlegen lässt, und protokolliert ihn in den Ergebnissen.

Engine activity as the liveness signal, ranks kept in step through autotune, OpenCode in the harness bench, catalog shapes in every record.

Added

  • ainode-bench harness --claude-effort <level> sets the reasoning effort Claude Code runs with (#127). Claude Code sends effort "high" by default, and a served chat template does not have to accept that value: Qwen3.8-Flash-Next takes only xhigh, medium and low, so every request came back API Error: 400 Unexpected reasoning effort high, the harness crashed in about 0.3 s per attempt and the model scored 0/10 on a suite the other three harnesses were passing. The same ten tasks then scored 8/10 and 10/10 at medium. The level is appended to the claude argv as --effort <level> and nothing else about the invocation moves; unset stays the default and sends nothing, so every number already recorded (Ornith, Qwen3.8 27B) keeps its meaning. It is written down where a reader will find it: the run's settings as claude_effort and the claude block's options as effort, both absent when the flag was not used, because a run with the agent's own default is a different statement from a null. Docs: the claude section of bench/harness/README.md.

Fixed …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.19: Anthropic Messages API über Fleet-Endpunkt

AINode 0.5.19 leitet die Anthropic Messages API (/v1/messages und /v1/messages/count_tokens) über den Fleet-Endpunkt mit Routing, Failover und Streaming weiter und erkennt Bild- und Dokumentblöcke dieser API.

The fleet endpoint forwards the Anthropic Messages API; Claude Code is a harness bench adapter; Qwen3.8-Flash-Next entry pinned and proven.

Fixed

  • The fleet endpoint forwards the Anthropic Messages API. vLLM serves POST /v1/messages natively alongside its OpenAI paths, but AINode's port 3000 registered only /v1/chat/completions and /v1/completions, so /v1/messages was a 404 on every node: a client that speaks Messages and nothing else (Claude Code, the Anthropic SDKs) had to be aimed at one engine's :8000 and lost the model-id lookup, capability-aware ordering and failover that every other protocol gets for free. /v1/messages and /v1/messages/count_tokens now go through proxy_to_vllm, the same handler as chat completions, so routing on the body's model, transport failover across every node serving it, the 404 for a model nobody serves, SSE streaming and header passthrough (x-api-key, anthropic-version, anthropic-beta reach the engine untouched) are the code that already existed rather than a second implementation. is_multimodal_request now also reads the Messages API's own spelling of media (an {"type": "image"} or {"type": "document"} block, including one nested in a tool_result), so an image sent over /v1/messages is ordered onto an instance that accepts images instead of being routed on the model id alone. The proxy also stops dropping the caller's query string (path_qs, not path): Claude Code posts to `/v1/messages?beta=true…

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.18: Verteilter Start nutzt heruntergeladenes Modell direkt

AINode 0.5.18 sorgt dafür, dass ein verteilter Start ein über AINode heruntergeladenes Modell direkt nutzt, statt es auf jedem Node erneut herunterzuladen.

The distributed launch serves a model downloaded through AINode; harness bench; Qwen3.8-Flash-Next entry.

Fixed

  • A distributed launch now serves a model that was downloaded through AINode instead of making every node re-download it. POST /api/models/download-repo writes a flat <models_dir>/<owner--name> directory, not the Hugging Face cache layout, and the solo launch has served that directory for releases. The distributed shapes did not: both served the repo id, the head and peer containers mounted only the HF cache, and _ensure_peer_has_model distributed only a hub/models--<owner>--<name> entry, so a 133 GB model already sitting on the head was pulled again on every rank, over the WAN, for every launch. Both shapes now go through one resolver: when that flat directory exists on the head and the mount is trustworthy (the same AINODE_HOST_HOME test the solo path uses), every rank serves /ainode-models/<owner--name> with --served-model-name <repo id> so /v1/models is unchanged, and every rank mounts its own copy of the store there: the head's models_dir, a peer's /home/<ssh_user>/ainode-nvidia-models (chosen like the peer HF cache, because we ssh in as that user). _ensure_peer_has_model ships the flat directory over the fabric (rsync when present, else tar over ssh, skipped when the peer already has it) instead of the hub entry, and the peer's mkdir -p covers the new path so docker can never invent a root-owned empty bind source. With no such…

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.17: Längeres Log-Stille-Fenster beim Warten auf Engine-Bindung

AINode 0.5.17 erhöht das Log-Stille-Fenster beim Warten auf das Binden der Engine von 360 s auf 900 s, damit lange Autotune-Pausen nicht mehr als Fehler gelten.

The bind wait tolerates a long autotune pause.

Changed

  • The bind wait's log-silence window defaults to 900 s, up from 360 s. Nemotron 3.5 Lightning on the GX10 goes quiet for 363 s during FlashInfer autotune and graph capture, so the 360 s window still declared a healthy start silent on the 0.5.16 roll of Spark-4 (never bound on :8000 after 870s (log silent for 363s); relaunching once). The relaunch is a no-op since 0.5.13, but the verdict spends the engine's single retry and the line misleads. Fifteen minutes covers every startup pause measured so far; the 1800 s ceiling still bounds a wedged engine.

Image: ghcr.io/getainode/ainode:0.5.17

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.16: Thinking-Schalter erreicht DeepSeek V4, Modellkarte korrigiert

AINode 0.5.16 lässt den Thinking-Schalter im Chat und im Bench nun DeepSeek V4 erreichen und behebt die abgeschnittene Modellkarte in der linken Chat-Leiste.

The chat model card fits its panel; the Thinking toggle reaches DeepSeek V4.

Fixed

  • The chat Thinking toggle and the bench's reasoning section now reach DeepSeek V4. Both sent only Qwen's enable_thinking switch, and the toggle sent nothing at all when ON, so a DeepSeek engine launched with thinking off by default could never be turned on from the UI, and the bench's DeepSeek "thinking on" measurement was thinking off. Both states are now sent explicitly under both template switch names (enable_thinking and thinking); a template ignores the name it does not read.
  • The chat MODEL CARD stays inside the left rail at any panel width. An old MODELS LIST rule still declared .model-card, and the chat rail's card is the only element left carrying that class, so it picked up align-items: center and padding: 20px. On a column flex container the cross axis is horizontal, so centring sized the card's head and body to their own content instead of the rail's width: with DeepSeek V4 Flash selected the body measured 368 px inside a 233 px card and overflow: hidden cut roughly 67 px off each side, clipping the leading characters of the model id, the section labels and the chips on the left and the values on the right. The dead rule is gone and the rail card now stretches its children, and the pieces that can hold an unbreakable token wrap instead of widening the box: the label and value grid uses minmax(0, ...) tracks, the id, HF link, engine image, capability no…

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.15: Log-Stille-Fenster auf 360 s erhöht

AINode 0.5.15 erhöht das Log-Stille-Fenster beim Warten auf das Binden der Engine von 120 s auf 360 s, damit langsames Laden der Gewichte nicht als Fehler gewertet wird.

The bind wait tolerates a real weight-loading pause.

Changed

  • The bind wait's log-silence window defaults to 360 s, up from 120 s. vLLM prints its weight-loading progress bar once per shard; a two-shard checkpoint (Qwen3.8 27B NVFP4 on Spark-1) is quiet for 206 s between updates, so the old window declared a healthy load silent, logged never bound ... (log silent for 121s); relaunching once, and spent the engine's single retry on a no-op. Six minutes covers the observed gaps with margin and the 1800 s ceiling still bounds a truly wedged engine. engine_bind_log_silence_seconds in config.json overrides it.

Image: ghcr.io/getainode/ainode:0.5.15

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.14: Solo-Engine-Logs nach losgelöstem Start verfolgt

AINode 0.5.14 verfolgt nach dem losgelösten Start die Logs der Solo-Engine, sodass das Warten auf das Binden die Ausgabe der Engine sieht.

The solo engine's logs are followed after the detached launch.

Fixed

  • The solo engine's logs are followed after the detached launch, so the bind wait sees the engine's own output instead of one container id and then silence. On the 0.5.13 roll of Spark-1 the wait declared Qwen3.8 silent after 121 s, moved on to Ornith while Qwen was still loading, and Ornith died with "No available memory for the cache blocks". start_solo now attaches docker logs -f to the solo log the moment the container is confirmed running, the same follower the mp head uses, so serialized replay actually waits for the primary to bind.

Image: ghcr.io/getainode/ainode:0.5.14

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.13: Lebendigkeit der Solo-Engine über den Container geprüft

AINode 0.5.13 beurteilt die Lebendigkeit einer Solo-Engine nun über den Container statt über den Start-Client, sodass sie nicht mehr fälschlich als beendet gilt.

Engine liveness comes from the container, not the launch client.

Fixed

  • A solo engine is no longer declared dead the moment its docker run -d client returns. 0.5.12 made the solo launch detached (#85), but the bind wait still judged death from the launch subprocess, so every solo launch read as container exited after 1 to 11 s and got a relaunch that failed on the container-name conflict with its own live engine (seen on every 0.5.12 roll of Spark-1). The backend now answers engine_exited() from docker inspect, is_running() asks docker for every launch shape, and the bind wait trusts that before the subprocess.

Image: ghcr.io/getainode/ainode:0.5.13

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.12: Multimodale Anfragen und parallele Fleet-Ansicht

AINode 0.5.12 leitet multimodale Anfragen an Instanzen weiter, die Bilder akzeptieren, und bringt serialisiertes Startup-Replay sowie echte Parallelität in der Fleet-Ansicht.

Serialized startup replay, capability-aware routing, real parallelism in the fleet view.

Fixed

  • A multimodal request is routed to an instance that accepts images (#83). Ornith 1.5 is served twice under one model id, text-only on Spark-1 (stacked, --limit-mm-per-prompt '{"image":0,"video":0}') and with vision on Spark-3, and the fleet proxy ordered candidates local-first on the model id alone, so a chat completion carrying an image_url landed on the text-only instance and came back as vLLM's At most 0 image(s) may be provided in one prompt even though another node could serve it. A request carrying an image_url, input_audio, video_url or file part now consults the capability cache /api/models/caps already fills: instances cached as accepting images go first, never-probed ones next, cached refusers are dropped, and that specific 400 is treated as a routing miss (record vision: false, fail over) instead of an answer. Every other 4xx still goes straight back to the caller, un-retried. When no instance of a loaded model will take an image, the 400 says so and names the nodes tried. One cache, not two: the master probes a peer's engine port directly, so its own cache already describes remote instances. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.11: NCCL_IB_HCA auf aktiven RoCE-Port beschränkt

AINode 0.5.11 beschränkt NCCL_IB_HCA auf den aktiven RoCE-Port hinter dem Cluster-Interface, sodass Nodes mit mehreren RoCE-NICs keine Verbindungs-Timeouts mehr erzeugen.

NCCL_IB_HCA names only the port behind the cluster interface.

Fixed

  • NCCL_IB_HCA names only the RoCE port behind the cluster interface (#84). The whitelist accepted every HCA with an IPv4 GID, so on a node with more than one RoCE NIC it also listed a live direct-connect port (Spark-2's rocep1s0f0 on 10.0.0.x) and a link-local autoconf port. NCCL pairs HCAs by list position across ranks, tried that port against the peer, and died in ibv_modify_qp with a connection timeout. The list is now filtered to the ACTIVE port whose GID address is the node's fabric IP, read straight from sysfs with no extra shell call, which is the port the proven recipes pin by hand.

Image: ghcr.io/getainode/ainode:0.5.11

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.10: /dev/infiniband in Engine-Containern eingebunden

AINode 0.5.10 bindet /dev/infiniband beim verteilten mp-Start auf einem echten Node in die Engine-Container ein, da die Erkennung nun über /sys/class/infiniband erfolgt.

The mp distributed launch maps RDMA devices on a real node.

Fixed

  • The mp distributed launch maps /dev/infiniband into the engine containers on a real node (#84). The presence check looked for /dev/infiniband from inside the AINode container, which sees sysfs but has no device node, so the head and worker were rendered with no --device and NCCL's IB plugin failed to initialise (NCCL_NET=IB, "Failed to initialize any NET plugin"). Presence is now judged from /sys/class/infiniband having entries, with the device path as the host-side fallback.

Image: ghcr.io/getainode/ainode:0.5.10

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.9: Verteilter Start ohne Ray und DeepSeek-V4-Flash-Rezept

AINode 0.5.9 führt mit der vLLM-mp-Multi-Node-Variante einen verteilten Start ohne Ray im Engine-Image ein und bringt Interface-Autodetect sowie das DeepSeek-V4-Flash-Rezept.

Distributed launch without Ray, interface autodetect, and the DeepSeek V4 Flash recipe.

Added

  • A distributed launch that needs no Ray in the engine image: the vLLM mp multi-node shape (#84). The distributed path could only do one shape, a ray start --head container plus SSH-launched ray start workers plus a docker exec of vllm serve --distributed-executor-backend ray inside the head, so it required the ray CLI in the engine image. Two engine images we actually need do not have it: stock vllm/vllm-openai:v0.27.1 (the head container exits 127, "ray: command not found") and the custom GB10 build that is the only thing serving DeepSeek V4 Flash correctly on sm121. NodeConfig gains distributed_executor (ray by default, or mp), and the mp shape runs one vllm serve container per node: rank 0 on the head, --node-rank k --headless on each peer over SSH, all rendezvousing on --master-addr/--master-port with vLLM's own multi-node executor. Peers go out before the head because the rendezvous, not a Ray port, is what waits. Container shape is the proven one: --network host --ipc host --shm-size 64g --ulimit memlock=-1 --ulimit stack=67108864 --gpus all, plus --device /dev/infiniband when the host has it. Serve args come from the same builder as every other launch, so a recipe flag still suppresses the built-in rather than duplicating it. Readiness, last_log_activity, the adaptive bind …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.8: Keine veralteten Assets, stabileres Startup-Replay

AINode 0.5.8 verhindert veraltete statische Assets nach Updates, zeigt die Modellkarte nur im Chat und lässt langsame, aber gesunde Engines beim Startup-Replay nicht mehr neu starten.

Fixed

  • Static assets can no longer go stale across an update (#77) — every /static URL in the served HTML carries ?v=<version> and /static responses are Cache-Control: no-cache. After the 0.5.7 roll a browser kept the 0.5.6 stylesheet on heuristic freshness and rendered the new chat unstyled.
  • The model card only renders on the Chat view (#77); it was showing on Cluster, Server, Models and Bench.
  • Startup replay no longer kills slow-but-healthy engines. The bind wait had a fixed 300s window. On vllm/vllm-openai:v0.27.1 a GB10 node spends minutes in FlashInfer fp4_gemm autotune and CUDA graph capture before the server listens, so on Spark-1 (2026-09-13) the window expired at 5 minutes on a 27B NVFP4 primary that bound at ~12 and a 35B-A3B stacked instance that bound at ~6. Both were relaunched from scratch, turning a 14-minute boot into 28. The wait now treats an engine as alive while its container is up and its log is still advancing, and relaunches only on evidence: container exited, log silent past engine_bind_log_silence_seconds (default 120), or engine_bind_ceiling_seconds reached (default 1800). An engine that dies on the way up still gets exactly one relaunch (0.5.5), and the ainode log now carries the reason and how long the wait lasted.

Changed

  • "Made in Texas" replaces "Powered by argentos.ai" in the web and onboarding footers, the CLI banner and status output, the bench report and …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.7: Flottenweiter Chat, Bench-Ansicht und Ornith 1.5 im Katalog

AINode 0.5.7 bringt einen flottenweiten Chat mit Modellkarte und Messwerten, eine Bench-Ansicht im Browser sowie Ornith 1.5 35B-A3B (NVFP4) im Katalog und behebt eine Race Condition beim Download-Fortschritt.

Added

  • Fleet-aware chat (#73) — the picker lists every ready instance across the cluster; a model card shows node, GPU, VRAM, TP, quant, engine image, speculative decoding and the engine's live context length; capabilities are probed with a real request (16 px image, dummy tool) rather than guessed; every turn carries measured TTFT, decode tok/s from server usage, reasoning tokens and the serving instance, with averages kept per instance. System prompt, temperature, max tokens, thinking toggle and stop.
  • Bench view (#75) — run the benchmark suite from the browser against any loaded instance: single stream, prefill scaling, sustained generation, concurrency, reasoning tax. Live progress and cancel, results table with JSON download in the public bench/SCHEMA.md shape, and the report rendered in-app. One run per node; refuses a not-ready instance; warns on a busy node. scripts/ainode-bench.py and bench/report.py are now thin shims over ainode/bench/.
  • Ornith 1.5 35B-A3B (NVFP4) in the curated catalog (#68), and bench/results/ with the measured runs behind the README table (#69, #70, #72).

Fixed

  • Download job progress race (#74) — the progress poller's last disk read could land after completion and overwrite 100% with a stale partial.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.6: Port-Prüfung, eigene extra_env und 64-MB-API-Limit

AINode 0.5.6 prüft bei der Port-Vergabe auf belegte Host-Ports, lässt Modelle eigene extra_env mitbringen, lädt Engine-Images mit Timeout vor und hebt das API-Anfragelimit von 1 MB auf konfigurierbare 64 MB an.

Port allocation now probes for host-held ports (a squatter on :8000 was silently killing engine launches on the master node). Models carry their own extra_env with recipe values beating the forced marlin default. Engine images pre-pull with a timeout instead of failing at launch. The API request ceiling moved from 1MB to a configurable 64MB so long-context requests stop bouncing.

Claude-Session: https://claude.ai/code/session_01YJ4KxqTc432yVpBBxR72zG

Co-authored-by: Claude Fable 5 noreply@anthropic.com

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.5: Einmaliger Neustart bei Engine-Absturz im Startup-Replay

AINode 0.5.5 startet eine Engine, die beim Hochfahren stirbt, im Startup-Replay einmalig neu, sobald die GPU wieder freigegeben ist.

Fixed

  • Startup replay retries an engine that dies on the way up (#63) — an engine can pass the launch check (its container reaches Running) and then die minutes later during weight load. Seen while rolling 0.5.4 onto a node: the startup sweep killed the running engine and replay relaunched immediately, while the driver was still releasing the GPU, so the nvidia hook handed the replacement no device (Can't initialize NVML, 0 active driver(s) found) and it exited during load. Nothing retried, so the node came back advertising nothing and the model stayed missing until someone re-loaded it by hand. _ensure_serving() now waits for each engine to bind and, if it never does, waits for the GPU to finish releasing and relaunches once — for both the boot primary and each replayed stacked instance. The retry is deliberately single: a model that fails twice has a real problem, and a loop would hide it.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.4: Eigene Engine-Images, neue Katalog-Rezepte, GB10-Korrekturen

AINode 0.5.4 lässt Modelle mit eigenem Engine-Image und eigenen Flags starten, ergänzt Katalog-Rezepte für Nemotron 3.5 Lightning und Qwen3.8-27B und korrigiert GB10-Workarounds sowie das vllm-serve-Argument-Handling.

Added

  • Models launch with their own engine image and flags (#62) — per-instance extra_vllm_args and engine_image on the node config, threaded through the per-load override set so they persist and survive restart-replay. A model whose published recipe needs flags AINode doesn't model (speculative decoding, MoE/mamba backends, reasoning + tool-call parsers) now launches through the normal load path instead of a hand-rolled container. A flag supplied by the caller suppresses the matching built-in rather than duplicating it.
  • Catalog recipes for NVIDIA Nemotron 3.5 Lightning 30B-A3B (NVFP4) and Qwen3.8-27B (NVFP4, native vision), each carrying its proven engine image, flag set, and recommended memory fraction — applied as defaults so a bare {"model": ...} load (what the dashboard sends) launches correctly.

Fixed

  • The 0.17-era GB10 workarounds (--enforce-eager, the NVFP4 MARLIN env) now apply only to the pinned default engine image. They are bugs in that build; on newer engines they only disable CUDA graphs and cost throughput.
  • vllm serve argv is normalized across engine images. vllm/vllm-openai bakes ENTRYPOINT ["vllm","serve"] while the default image uses NVIDIA's passthrough shim, so emitting our own prefix produced vllm serve vllm serve <model> and the engine exited with "unrecognized arguments". The entrypoint is deliberately not overridden — that would bypass CUDA setup. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

AINode

AINode 0.5.3: Wahrheitsgemäße Instanzen, flottenweites Update-Banner

AINode 0.5.3 zeigt Instanzen wahrheitsgemäß mit Node-Namen und korrektem Status an, und das Update-Banner aktualisiert nun die gesamte Flotte statt nur den Head-Node.

[0.5.3] — 2026-07-07

Fixed

  • Truthful instances everywhere (#59) — every instance card shows its node NAME (not a hex id); the Server view lists stacked instances with node + port and its count matches reality (Eject only where it can actually target); per-instance status probes both directions (starting → serving, and back to failed when an engine dies — no more stale STARTING bars or false READY); the top-bar update banner now updates the whole fleet (/api/cluster/update-all) with an honest confirm, not just the head node.

Originalquelle(öffnet in neuem Tab)Problem melden