Release notes
Eldric 5.0 is the current platform line. Patch releases ship regularly; the current build is always the one published as *-latest on repo.eldric.ai, which is what dnf installs and updates to. All packages and models are linked from the downloads page. This page summarizes what the 5.0 line delivers and what to check before you install or update.
What Eldric 5.0 delivers
Eldric 5.0 runs as a single platform you install and operate on your own infrastructure. Available in the current line:
- One package you install with
dnf(CPU), plus an optional CUDA build for GPU hosts, an ARM64 minimal build and a macOS package — all from repo.eldric.ai. - Multi-node clustering with role-based nodes, load-balanced routing and controller high-availability.
- An OpenAI-compatible public API through the Edge, secured with API keys — see the API reference.
- Retrieval-augmented generation with on-device embeddings and per-tenant knowledge bases — see Using RAG.
- Multi-agent orchestration and multi-tenant isolation.
- A unified registry of scientific data sources, plus speech-to-text / text-to-speech, messaging and IoT capabilities.
- Native GGUF inference without an external runtime.
The 5.0.x update line
The 5.0.x line ships stability, security and bug-fix updates on top of GA, plus incremental capabilities delivered as patches — including controller high-availability, multi-site federation and opt-in smart-memory inference. The latest packages are always published as *-latest on repo.eldric.ai; update with dnf.
CPU and GPU are separate channels
The CPU package and the CUDA package for GPU hosts are built and published independently, so they are not always at the same patch version. A mixed cluster can legitimately be running one version on its CPU nodes and another on its GPU nodes, and dnf update on each host gives you whatever is current for that package — not a single cluster-wide version.
Check what each channel currently offers before assuming they match:
dnf --showduplicates list eldric-aios
dnf --showduplicates list eldric-aios-cuda
If you need both at a specific version, pin the version explicitly rather than relying on *-latest to line up. Every node also reports its own running version at /health, which is the honest answer for a given host. From 5.0.168 that path needs a login from outside, so ask on the host itself: curl http://localhost:8880/health.
5.0.177 — security release
Published for the CPU package and the GPU package, for Fedora 43 and Fedora 44 on x86_64 and aarch64 (GPU: Fedora 43 on x86_64). 5.0.176 was not published; 5.0.177 follows 5.0.175 and contains everything listed here. Update all nodes of a cluster, and update the edge nodes and the controllers first: 5.0.177 changes how a forwarding node passes on who the caller is, and an older forwarding node in front of updated nodes can lose that information. While nodes run different versions, a knowledge-base search that reaches an older data node can miss that node's results; they are reported as an error, not silently.
⚠ Security notice
| What could happen | Affected releases |
|---|---|
| Requests that passed through an edge or another forwarding node could run with node rights. A signed-in user, a viewer included, could through such a node perform administrator actions on the controller and on other nodes. A forwarding node now passes on the caller's identity, and the request runs with the caller's own rights. | 5.0.0 – 5.0.175 |
| Backups could be listed, created, restored and redirected by any signed-in user, including setting where a backup is sent; an expired or invalid licence did not stop backup operations either. Backups now need an administrator. | 5.0.0 – 5.0.175 |
| 16 controller actions that change state did not check the caller's role, among them paid model work, dream cycles over all users' sessions, namespace changes and federated training jobs. They now need an administrator. | 5.0.0 – 5.0.175, depending on the action |
| Further actions in other modules lacked a role check: IoT writes to devices and control loops; reading server files through media tools; registering image providers; storing into or loading other tenants' memory files; data sources that could run SQL on server files, connectors and memory settings; training, run approvals, worker registration and global skills; and a knowledge-base search in the chat that could reach another tenant's documents. They now need an administrator or are limited to the caller's tenant. | 5.0.0 – 5.0.175, depending on the action |
| A worker could be registered without a login through the edge. | 5.0.0 – 5.0.175 |
| A node-to-node knowledge-base search accepted callers without a login and returned document text of any tenant. It now accepts only cluster nodes. | 5.0.0 – 5.0.175 |
Sign-in and password reset had no rate limit. They are now limited per client address and per account; a refusal answers 429 too_many_attempts with Retry-After. | 5.0.0 – 5.0.175 |
| Revoking a capability token did not revoke tokens derived from it. It now does, for tokens derived after the update. | 5.0.0 – 5.0.175 |
| A forgotten memory could come back. Forgetting did not survive a restart or a switch to the standby; after a restart a forget did not reach memories loaded from disk; the text of a forgotten entry could stay in the memory's metadata, which feeds training extraction; and a forget of a private memory looked in the wrong place. A forget now also reaches the standby data node at once. | 5.0.0 – 5.0.175 |
| An account that existed only on a data or edge node could be taken into the cluster's accounts, with its role, including an administrator. The controller now takes accounts only from other controllers. | 5.0.12 – 5.0.175 |
| A signup on an edge node created a local account there, valid only on that node; on a node with no accounts yet (for example a newly added edge or data node) the first such signup became a super-administrator there. Signup, login and password reset on every node now go to the controller. | 5.0.0 – 5.0.175 |
After updating:
- Review settings that any signed-in user could change before; changes made then are still in place: backup schedules and their destinations; registered workers and swarm agent-workers; image providers (address and API key); data sources, connectors, memory settings and cluster file rules; the telemetry (OTLP) export endpoint, through which data leaves the cluster; tenants, namespaces and cluster alerts; skills marked as global; the router strategy; IoT device bindings.
- Accounts: check the controller's account list for administrator or super-administrator accounts you do not recognise, especially ones created around the time an edge or data node was added. Also check
GET /api/v1/identity/local-unknownon each data and edge node, for local accounts the controller never took over. - Capability tokens: a token derived from another one before the update does not know its parent, so revoking the parent does not revoke it. Revoke such tokens individually, or let them expire.
- Forget and erasure requests: if you used forget for an erasure request on 5.0.175 or earlier, repeat it after updating. The repeated forget also removes what the earlier one left behind (the key and value text in the memory's metadata, and the memory's association in the private store), also after a restart, and reaches a standby data node within seconds.
New and changed
- Log collection and security analytics (new): on a node you give the
logsrole, Eldric indexes what rsyslog collects from your servers, switches, firewalls and other devices, writes and checks rsyslog's configuration itself, keeps two copies on two disks, removes secrets before indexing, and keeps logs per class for as long as you set. Administrators search and follow logs live in the admin console's Syslog dashboard and in the chat, and get alerts from Sigma detection rules, the Spamhaus DROP list, learned baselines and incidents; NetFlow/IPFIX volumes are kept as aggregates. No licence feature is needed. See Syslog. - Syslog forwarding settings are saved: a setting made with
PUT /api/v1/system/syslog/confignow survives a restart and wins over theELDRIC_SYSLOG_*environment;DELETEon the same path returns to the environment, andGETshowssource. Invalid values are refused with400instead of being silently replaced. - Accounts on every node are answered by the controller. If no controller can be reached, signup and login answer
503 account_service_unreachable; no local account is created instead. The chat then says the account service is unreachable, not “wrong password”, and does not log you out. - High availability for data nodes: memory and tenant files reach the standby data node every two minutes by default, or on demand; models are mirrored in full, and a model deleted on the primary is kept on the standby. A standby is promoted only when it has caught up; a manual promotion over a standby that is behind answers
409 replica_behindunless forced. A failed primary that comes back can no longer overwrite its replacement. Automatic failover is off by default; turn it on only after testing on your own nodes. Known issue: knowledge-base content that existed before replication was switched on does not reach the standby, while the standby is still reported as caught up; see Known issues before you rely on a standby. - Raspberry Pi: the aarch64 package is tuned for the Raspberry Pi 5 and CPUs with the same instructions. On a Raspberry Pi 4, local models are not available: those requests answer
503 cpu_lacks_arm_extensionswith the reason, and the rest of the node keeps running. - Chat:
/chat#ask=<question>opens the chat with the question filled in, without sending it.
5.0.175 — security release
Published for the CPU package and the GPU package. Update all nodes of a cluster. 5.0.175 contains three security fixes and new features.
⚠ Security notice
| What could happen | Affected releases |
|---|---|
| A cluster node without its own user accounts answered every request without authentication. A data, inference or edge node that had joined a cluster but held no user accounts of its own ran with authentication switched off: anyone who could reach its port could read and change data on the node, including administrative settings such as where the node replicates its data to. Controllers were protected from 5.0.12 on, once their first user account existed. A node now enforces authentication whenever it is a cluster member, and checks keys with the controller. | 5.0.0 – 5.0.174 |
| A deleted knowledge-base document could still be found by search. Deleting a document answered success but left its text in the index that search reads; “delete, then re-ingest” left two copies. Bulk deletes had the same gap. | 5.0.0 – 5.0.174 |
| The chat's forget tool could delete other users' memories in bulk again. The 5.0.172 fix was missing in 5.0.173 and 5.0.174: the chat assistant's forget tool could be made to clear a whole memory namespace for an ordinary user, including other users' private memories, also through instructions hidden in a document or web page the model reads. The direct memory API was not affected. | 5.0.173 – 5.0.174 |
After updating:
- Open nodes: update every node of a cluster. If a data, inference or edge node of yours was reachable from untrusted networks, treat its data as possibly read or changed by others, and check where it replicates to. Past access cannot be reconstructed, because the node did not authenticate those calls.
- Deleted documents: an administrator can list what earlier deletes left behind, per knowledge base, with
GET /api/v1/vector/namespaces/<tenant>/<namespace>/shard-only(read-only; each entry is markeddeleted_likelyorundecidable), and remove each entry with the normalDELETE /api/v1/vector/documents/<tenant>/<namespace>/<id>, which on 5.0.175 removes it from search. With a standby data node configured, the delete reaches the standby too. If you deleted documents to meet an erasure request, check those knowledge bases.
New and changed
- Refused credentials answer 401: a key or token the node does not accept now answers
401 invalid_credentialwith areasoninstead of403. If no controller can be reached to check a key, the answer is503 credential_check_unavailable: the key was not judged, so keep it. See the API reference. The chat returns to the login when its key is revoked. - User accounts and key changes reach every node within one heartbeat.
- The CPU package runs local models itself (embeddings and chat with GGUF models), without Ollama or another backend. On x86_64 it needs a CPU with AVX2; without it, the node reports why in
/health, those requests answer503 cpu_lacks_avx2, and the rest of the node keeps running. On a single server without the cloud module,POST /v1/embeddingsuses this local inference. - Knowledge-base writes are replicated to the configured replication peers, synchronously or in the background, with a backlog that catches up after an outage; deletes and namespace changes replicate too. Writes and uploads go to a live data node, and upload jobs retry and resume. Upload bytes are no longer stored on a node without the data role.
- Settings that lived on one controller are replicated to every controller: roles and user grants, the storage policy, webhook subscriptions, dream settings and knowledge-base settings. Before, a controller failover lost them.
- Switching the data role off on a node that still holds data is refused (
409 data_role_holds_data, listing what would be left behind) unless forced;/healthwarns when no node runs the data role. - A backup that could not include a requested part is reported as partial, naming what is missing, instead of “success”.
- A failing background task no longer stops the whole node; it is logged and restarted.
- Knowledge bases: each chunk records the embedding model that actually produced it, and an administrator can relabel a knowledge base's embedding model when the stored vectors prove the new label correct.
- Router: the packaged router model now takes precedence over router models placed by hand in the data directory. To keep your own, pin it in the admin interface or with
ELDRIC_ROUTER_XLSTM_MODEL; the admin pin wins.GET /api/v1/router-xlstm/statusreports which file was loaded and why (source). - Science: 17 sources that could not be reached through
/api/v1/science/querynow can, and results carryrecords[]in one common format beside each source's raw answer. - Plugins: new VMware ESXi plugin for a single host (inventory and state; creating, changing and deleting VMs and snapshots only when registered for writing). FortiGate: configuration tools and interface history, and a read error is no longer reported as “no events”. Note: a FortiGate API token you consider read-only may still accept writes; the plugin's read-only registration is what stops it, so also give the token a read-only admin profile on the FortiGate. Netatmo: heating per room, heating schedules, history, Legrand/BTicino devices.
- Packages: each Fedora release gets its own build: Fedora 43 and Fedora 44, on x86_64 and aarch64. The GPU package is built for Fedora 43 on x86_64.
5.0.174
Published for both the CPU package and the CUDA package. Update all nodes of a cluster. 5.0.174 contains a security fix, a fix for a data-loss defect, and new features.
⚠ Security notice
| What could happen | Affected releases |
|---|---|
Swarms, their runs, goals and messages were not limited to their own tenant. A signed-in user of another tenant who knew or listed an id could read a swarm (including its goal text and the input prompts of its runs), run, abort or delete it, approve or abort its goals, and read its messages. Records of another tenant now answer 404, exactly like an unknown id. | up to 5.0.173 |
After updating: run records written before 5.0.174 carry no tenant and are visible to administrators only.
⚠ Data loss: world-model files loaded through two administrator routes
A memory file (.emm) that contains a world model lost its world-model part when it was loaded through either of two administrator routes: POST /api/v1/worldmodel/rollout with emm_path (the file was rewritten on the same call), and POST /api/v1/retrieve/load (the file was rewritten when the module or the node stopped). Memory files without a world model were not affected. Affected: 5.0.17 up to and including 5.0.173.
- Until you have updated to 5.0.174, do not use either route on a file that contains a world model. On older releases every such use removes the world-model part.
- If an older release answered
no_world_modelfor a file you know is a world model, that answer was wrong: the file still had its world model until that same call removed it. - After updating to 5.0.174, check each world-model file you used with either route by running one rollout with it. From 5.0.174 the rollout opens the file read-only, so the check changes nothing. Do not run this check on an older release; there it is exactly the call that removes the world model.
- A normal rollout result: the file is intact.
no_world_modelfor a file you know is a world model: its world-model part is gone or damaged. Restore the file from a backup taken before you first used it with either route. The update cannot restore a part that is already lost.no_world_modelfor a file that is only memory, without a world model, is expected and means nothing was lost.nslot_load_failed: the world-model part is present but cannot be read as a model. That is not this data loss; it points to a different problem with the file or a format this release does not know. Contact support before restoring a backup.
- If you never used these routes on world-model files, nothing is affected. Eldric's own dashboards, chat and router do not call either route.
New and changed
- API keys in the chat: user menu → API keys → Generate a new key…. The new key is shown once; the chat switches to it immediately, and the old key stops working everywhere. Update every script and app that used the old key. Administrators can reset any user's key on the users page (Reset API key).
- A cloud provider's refusal reaches you instead of a generic
model_not_ready: for example missing credit (402), invalid credentials (401/403), or a model the provider does not offer. - Re-check an excluded cloud backend: after topping up an account, an administrator presses Re-check in the cloud dashboard (or
POST /api/v1/cloud/backends/<id>/recheck). It makes one minimal request to the provider; no restart is needed. - The model list says why a cloud model is unavailable:
serving_state_reasonisprovider_rejected(with the provider's status) orunreachable, and the chat's model picker names the reason, including models the provider no longer lists and models an administrator listed that the provider has not confirmed. - The router's tool-selection model now ships in the packages, so tool preselection works on a fresh install without a separate download. Its embedding model's licence notice (bge-m3, MIT) is included.
POST /api/v1/exmm/consolidateis fixed (known issue in 5.0.173): every memory section is kept, the file is backed up as<file>.pre-consolidate.bakbefore it is overwritten, and a file the operation cannot handle is refused with a reason.- Memory files are written only when they changed. Files opened only for reading (world-model rollout with
emm_path, registering an.emmin the model store) are never written, so they no longer get a new modification time when the service stops. Do not use the modification time as a "last used" signal.
5.0.173
Published for both the CPU package and the CUDA package. Update all nodes of a cluster. 5.0.173 contains security fixes, a one-time migration of the memory database, and new features.
⚠ Security notice
| What could happen | Affected releases |
|---|---|
| Any signed-in user could read and overwrite edge plugin settings, which often hold API keys. Only values an administrator had stored through the plugin-settings API were exposed; before 5.0.173 those values were never handed to the plugin. | up to 5.0.172 |
| One user could overwrite or forget another user's private memory by storing or forgetting the same key in the same namespace. An overwritten entry was gone after the next restart. | 5.0.16 – 5.0.172 |
| A chat tool hidden from a role could still be run by name: a tool's minimum role was only used to decide which tools to offer, not checked when one ran. | up to 5.0.172 |
| Several routes trusted the tenant named in the request (URL ingest, workspaces, data graph, leases, name registry, agent chat, brain chat, swarm tasks), and knowledge-base wizard sessions had guessable ids without an owner check. | up to 5.0.172 |
| Upload progress showed any upload's file name and sizes to anyone who knew its id. | up to 5.0.172 |
| Model routes opened any file path given in a request. | up to 5.0.172 |
| Extension data, including stored secrets, was readable by other local users. | up to 5.0.172 |
After updating: rotate any API key you had stored in plugin settings. Memory entries that were already overwritten cannot be recovered. A model listed in exmm_preload by an absolute path outside the model directory is now skipped with a warning: move it into the model directory or load it by name.
Memory: one-time migration
On the first start, the memory database is migrated so that each user's entries are kept apart. A backup is written first, to <memory.db>.pre-slot-pk.bak, and the migration is committed only if the row counts match. To roll back, stop the node, move memory.db and its -wal/-shm files aside, and restore the backup as memory.db; entries written after the update are lost.
New
- Rotate API keys.
POST /api/v1/users/me/api-keyfor your own key, andPOST /api/v1/identity/users/<id>/api-keyfor administrators. The new key is shown once; the old key stops working on every node at once. - Memory recall ranks by relevance when given a
query; each hit has ascore, andlimit/top_kapplies after ranking. - Memory wizard: turn a document into memory entries, private by default, with settings suggested by document type and a way to forget a document's entries again. In the chat:
/memorize. - Large files are indexed. PDF and Office files over 4 MB are indexed as a background job with progress. Latin-1 and UTF-16 text is indexed too. A single request over 128 MB answers
413and points to the chunked upload. Still not indexed: scans without a text layer, and.doc,.xls,.rtf. - Forecasting: streams, embeddings and trained heads, covariates, context padding and chosen quantiles. See help, topic "Forecast streams". A model that needs a fixed context length now refuses any other length instead of truncating it silently. Renamed: the load error
forecast_model_load_failed(the previous code stays indeprecated_errorfor one release) and the backend idexmm-patch-forecaster. - Plugin settings are managed by administrators, with secrets masked and values checked against the plugin's schema. Server plugins can be switched on and off (
POST /api/v1/edge/plugins/<id>/enable/disable). - Six edge plugins ship in the packages (audit-logger, markdown-export, research-graph, system-prompt-manager, token-counter, web-search). They are brought into the node at each start without touching your settings or plugins you switched off.
- Unhealthy backends are skipped: their models show
serving_state: "backend_unhealthy", and requests go elsewhere. - Provider keys without a backend are reported to administrators; key values are never shown.
- Faster, clean shutdown, within seconds.
- Third-party notices ship with every package, under
/usr/share/licenses/<package>/. - WhatsApp bridge: the package no longer bundles its Node.js modules. To use it, install npm and run
cd /usr/share/eldric/whatsapp-bridge && npm install. Those modules include libsignal, licensed under GPL-3.0; you install it yourself, Eldric does not distribute it.
5.0.172 — security release
Published for both the CPU package and the CUDA package. Update to 5.0.172. It closes several defects in memory and dreaming, some of them present in releases you may be running.
⚠ Security notice
| What could happen | Affected releases |
|---|---|
| Private memories were copied into the cluster memory. While an administrator had pinned a master memory, every stored memory entry, private ones and dream themes included, was copied into the master's system store. These copies cannot be read back through memory recall. From 5.0.107 they could also be folded into the pinned model file; other tenants are only affected if that happened and the file is loaded or served. | 5.0.62 – 5.0.171, only with a pinned master |
| A signed-in user could write into another user's private memory through internal write targets meant for replication between nodes. | up to 5.0.171 |
| Other users' memories could be deleted in bulk, in two ways: through the memory API by any signed-in user, and through the chat's forget tool, which passed the model's arguments on unchanged. The second way could be triggered by instructions hidden in a document or web page the model reads (prompt injection). The fix for the forget tool is missing in 5.0.173 and 5.0.174 and back in 5.0.175; see the regression note under 5.0.174. | 5.0.0 – 5.0.171 (private memories exist from 5.0.16) |
| A viewer could start a dream cycle for another user's scope. | up to 5.0.171 |
Clean-up if you had a master memory pinned
- Check whether a master is pinned:
GET /api/v1/cluster/storage/system-exmm. - On the node holding the pin, and on every copy, delete the system-store copies
<memory>/_system/<model_id>_*. - Check the model file for such sections with
GET /api/v1/exmm/models. If present, remove them withDELETE /api/v1/exmm/shard. - If sections were already folded into the base, rebuild the model file from a backup taken before the fold.
- Keep the files
u:usr-…_dreams_*: those are each user's own dream themes. - If you no longer need it, remove the master pin.
Changes
- Only cluster-scoped memory reaches pinned memory containers. From 5.0.172 an entry is written into a pinned container (master, tier or project) only when it was stored with scope
clusterorpublic, which already needsmemory:write_publicor superadmin. Private, shared, group, project and tenant entries stay in their own scope, and dream themes no longer feed the master. To put knowledge into the master, store it with scopecluster. - Bulk deletion of memories needs an administrator, through the API and through the chat's forget tool.
- Internal replication targets accept only authenticated nodes.
- Dream cycle and dream configuration need an administrator. A wrongly typed value now answers
400instead of500. - Distill keeps its file paths inside the data directory on
/api/v1/distill/documentand/api/v1/distill/jobs, as/api/v1/distill/runalready did. - Users without a tenant can use agent features. This fixes the 5.0.171 known issue for, among others, the first administrator.
- Tenant suspension and deletion take effect across the whole cluster.
- The chat says when an uploaded file is stored but not indexed, and names the reason, instead of counting it as indexed.
- The chat's memory commands work, or say why not. "Remember as …",
/rememberand/forgetnow send what the server expects. A refusal, for example because memory is not in your licence, is shown as Not remembered — reason instead of a false success.
5.0.171
Published for both the CPU package and the CUDA package.
- Chat with models served by another node works again. In 5.0.170, on a node that runs both a router and a cloud module, a chat request for a model that only another node serves failed with
502 model_not_ready. 5.0.171 forwards such requests to the node that serves the model, as 5.0.169 did.
5.0.170
Published for both the CPU package and the CUDA package. Update all nodes of a cluster together: a 5.0.169 node calling a 5.0.170 node gets 401 on the embedding fallback.
Breaking changes — prepare before you update
- Chat, completions and embeddings always need a login, inside your network too. Once authentication is on,
/v1/chat/completions,/v1/completions,/v1/embeddings,/v1/cloud/embeddingsand/v1/inference/embeddingsanswer401to any call without an API key or session, including calls from your LAN or straight to port 8880. The reverse-proxy header no longer matters for these paths./v1/modelsstays open. Give every integration that calls them anX-API-Keyor aBearersession first. - Eldric's own tools on the edge become opt-in. On
POST /v1/chat/completionsthe edge runs Eldric's tools (web, files, knowledge and others) only when the request carriesX-Eldric-Edge-Tools: 1. Eldric's web chat and the macOS, iOS and CLI clients send it. Without it, the request passes through with standard OpenAI behaviour: the client's own tools reach the model,tool_callscome back to the client, and the edge runs nothing. Every response says which mode applied (X-Eldric-Edge-Tools: onorpassthrough). OpenAI-compatible integrations such as OpenWebUI that relied on Eldric's tools without sending their own should send the header.
Security hardening
- The dream engine now attaches the node credential only to genuine local addresses. Before, an administrator-configured engine URL that only looked local could receive it. This concerns releases from 5.0.49 to 5.0.169. If you have changed the dream engine's LLM URL, check that it points where you intend.
- Only real loopback addresses count as local. Short forms such as
127.1in a local LLM URL no longer count, and calls there answer401. Use127.0.0.1orlocalhost.http://[::1]now counts as local.
Other changes
- An invalid API key on
/v1/*answers401, withWWW-Authenticate: Bearerand codeinvalid_api_key, as OpenAI and Anthropic SDKs expect./api/v1/*keeps answering403. - Models without tool support are never swapped for another family. When a request needs tools, only a tool-capable model of the same family may step in. Otherwise the answer is
422 model_cannot_tool_callwithsupported_models. - The router always answers chat where it runs. On nodes with both a router and a cloud module, knowledge-base injection, intent detection, telemetry and the
X-Eldric-Routeheader now apply too. - Agents reach the model through the cluster. On nodes without a local model, agents used to reply
(inference unavailable)as a success. They now reach the serving node, or return the errorno_model_answered. - The public model list names the serving node instead of
self, through an edge without its own cloud module. POST /api/v1/tools/executeon the edge acceptsargumentsas well asargs.- The agent's
default_modelsetting (andELDRIC_AGENT_DEFAULT_MODEL) now takes effect. - The controller no longer reads its node credential while another thread renews it.
- Superadmin is granted only on self-registration, and once per cluster. Before, the first human user of each node became superadmin by any path, so a viewer created by an administrator could become superadmin. Now only a self-registration can be promoted, and only if it is the first in the whole cluster. If a node cannot reach the controller to check, nobody is promoted, and an administrator can raise the role later.
5.0.169
Published for both the CPU package and the CUDA package.
Clustering and routing
- Nodes learn from the controller which modules run where. Every registration and heartbeat (every 30 s) now returns a module map: for each node, its address, node id, running modules and when it was last seen. Nodes merge the map rather than replace it. An entry seen in the last 90 s is kept even if a fresh map lacks it, so a newly elected leader does not make live nodes disappear. The map is saved to
<data dir>/cluster/module-map.jsonand loaded at start./healthreports its size, age and whether it is stale. Right after a controller failover the map is incomplete for about 30 s, until nodes have heartbeated to the new leader. - The chat model catalogue comes from the module map. A node asks every node that runs a cloud or inference module directly. Before, nodes without a controller fell back to the address list in
ELDRIC_AIOS_PEERS, so models served elsewhere could be missing and answered with422.ELDRIC_AIOS_PEERSis now only a bootstrap fallback until the first map arrives: cluster membership needs no environment file. - A router node dispatches chat itself. Before, a router without a cloud module answered
503 cloud dispatcher not loaded, even for models served natively elsewhere. It now uses local inference, otherwise the node that serves the model, otherwise an honest404. The model checks run once, in the same order on router and cloud nodes. Chat calls between nodes carry node authentication and a hop marker, so a request is never passed on in a loop. - Models this install cannot run are no longer offered. A model without a runtime on this installation is marked
no_runtime./v1/modelsno longer lists it, and a request for it answers422. Administrators still see it, with its state.
Reliability
- Nodes keep working after the node credential is renewed. The credential is renewed after half its 24-hour lifetime. The data module kept its first copy, so other nodes rejected its calls after the renewal. Every read now takes the current credential.
Security
- Advisory red-line check for administrators. New admin-only route
POST /api/v1/security/redline/check. Send anaction_description; it checks the action against the Guardian red lines, returns the verdict and writes it to the audit log. It is advisory and blocks nothing. Other callers get403.
The breaking changes that followed this release are listed under 5.0.170 above.
5.0.168
Published for both the CPU package and the CUDA package.
Security
- Inference and operational paths need a login from outside. Through the public edge,
/v1/chat/completionsand/v1/embeddingsnow answer401without credentials. So do/health,/metrics,/api/v1/system/version,/api/v1/system/incidents,/api/v1/license/healthand/api/v1/cloud/backends/public. Earlier releases answered these anonymously. Calls inside the cluster and on the host itself (curl http://localhost:8880/health) are unchanged, and clients that already send credentials are unaffected. - Action needed if you run your own reverse proxy. The service treats a request as coming from outside only when the proxy marks it with
X-Eldric-Public-Edge: 1. Without that header these paths stay open, exactly as before. In everylocationthat proxies to port 8880, set the header to1for public clients and0for your private ranges, and always overwrite whatever value the client sends. - Agent run, stop and rollback check who owns the agent.
POST /api/v1/agents/<id>/run,/stopand/rollbacknow refuse a caller from another tenant with403 cross_tenant_access_denied. Before, anyone signed in who knew an agent's id could start another tenant's agent, stop its executions or reset its runtime version. - The public model list no longer shows internal addresses. Through the edge,
GET /v1/modelsandGET /v1/cloud/modelsshow a node name in place of every private IP address, or[redacted]if the name is unknown. The fields themselves are kept, so clients that read them keep working.
Routing
- A model is only ever substituted by one of the same family. If the requested model is unavailable, Eldric tries the same family and size, then the same family in any size. If neither exists, the request fails with
422 model_not_available, listing the available models. It is no longer silently answered by an unrelated model, including for requests that need tools. A model name without a size suffix counts as its own family up to the first:. Sokimi-k2-thinkingis not replaced bykimi-k2. - Chat goes to the node that actually serves the model. The model catalogue used to record the node it asked rather than the node that answered, so a chat could be sent to a node with no chat route and fail with a 404. Forwarded answers now name the node that answered, and the catalogue uses that node.
Operations
- Controller VIP files ship in the packages.
/usr/libexec/eldric-aios/eldric-vip-check.shis a keepalived health check that succeeds only on the current Raft leader, so the virtual IP follows the leader./usr/share/eldric/keepalived/keepalived.conf.templateis a template to adapt. The package activates nothing: it writes nothing to/etc/keepalived/and starts no service. The VIP moves to a new leader, but modules that are not already active on that node are not started by the move.
Before updating
- Read Known issues for current rough areas and preview features.
- Confirm package availability for your platform on repo.eldric.ai.
- Test critical API calls against a staging tenant.
- Back up tenant and knowledge-base state using supported workflows.