mirror of
https://github.com/EasyTier/EasyTier.git
synced 2026-10-08 19:06:14 -08:00
fix(web): harden managed config sync between console and clients (#2567)
* fix(web): fence managed config runtime reconciliation Keep runtime reconciliation tied to the currently authorized session so stale connections cannot mutate a replacement session runtime. Accumulate only contiguous dirty IDs and load their latest SQLite state. Require the applied revision to match the earliest Patch base and the persisted revision to match the latest target. Otherwise, reconcile the full desired state. Use separate runtime-state and config-cache epochs. Managed updates can reuse observed configs; direct mutations invalidate them. Update sync documentation to match. * fix(web): interrupt validation retry on state changes Track meaningful validation state changes separately from periodic dirty signals. Applied revision changes wake a failed validation immediately, while heartbeat-driven revalidation retains the retry backoff. Treat Notify as a wake-up hint and recheck the state-change epoch after every wake so stored permits and periodic heartbeats cannot cause retry storms. * fix(web): retry unconfirmed connected webhooks Retry node-connected webhook delivery on retryable errors with a short 100ms/500ms backoff and give up immediately on non-retryable errors. Re-check that the session still owns the connection before every attempt and before recording the delivery, so a replaced session can no longer record a stale connected binding. * fix(web): fence disconnects by session ownership Return whether session removal actually removed the current route owner, and emit disconnected only for that owner. Replaced sessions can no longer invalidate a newer connected route. * fix(web): hot-patch managed hostnames Include hostname changes in the hot-patch path instead of falling back to a full restart. When a full overwrite run is required and the desired config has no hostname, inherit the current runtime hostname so an unmanaged value survives until it is explicitly cleared. Read back the runtime config after an overwrite run and verify it converged instead of assuming the desired state was applied. * fix(web): retry transient runtime reconciliation failures Keep the per-session managed runtime reconciliation worker alive when a single database round fails. Retry from the next heartbeat so persisted managed revisions can still converge after restart-time contention. Reserve terminal worker shutdown for destroyed session or storage state, and cover recovery after a transient revision read failure. * fix(web): accept omitted hostname after runtime apply Release 2.6.4 omits hostname from config readback when it matches the device hostname. Trust a successful hostname mutation only when the returned field is absent, while continuing to verify every other field and rejecting explicit mismatches. * fix(web): ignore unmanaged runtime device names Windows release 2.6.4 generates a random interface name when the managed config leaves dev_name empty. Exclude that runtime-owned value from reconciliation unless the desired config explicitly sets a non-empty device name, preventing endless overwrite restarts. * feat(web): report failed network instances to console Expose stopped Core instances with startup errors in heartbeats. Merge Core failures with direct managed-run RPC failures in easytier-web. Send failed instance IDs during token validation without error text. Prune local run failures when managed configs are deleted. * fix(web): distinguish unknown runtime application state Track whether the current session has observed its applied revision separately from the optional revision value. Report this fact through validate-token so Console can preserve application state across receiver restarts while recognizing deliberate pending mutations. * feat(web): configure heartbeat timing from server Heartbeat responses now provide the interval and RPC timeout. Legacy servers use local defaults and remote values are clamped. Web configuration and session receive timeout follow the policy. * fix(web): reject inactive control sessions Route control RPCs by machine id only to sessions whose RPC manager is still running, so a session that has been stopped or replaced can no longer receive control traffic addressed to the device. * fix(core): filter network info before collection When a collect-network-info request names specific instances, collect those instances only instead of collecting every instance and filtering the result afterwards, so unrequested instances no longer run per-collection work on every request. * feat(web): enable focused runtime diagnostics Enable easytier-web info logs by default while preserving explicit log configuration. Record startup settings, session lifecycle, failed instance changes, webhook queue and request latency, and managed runtime operation timings for production diagnosis. * fix(web): preserve managed revision across reconnects Keep one runtime identifier for each Core WebClient lifetime. Reuse its managed runtime state after transport reconnects. Retain applied revisions and reconcile hints while disconnected. Preserve runtime epochs so stale work cannot mark a revision applied. Reject stale sessions from reclaiming routes after reconnect. Core or Web restarts and legacy clients still use unknown state. Immediately revalidate a restored revision after authentication. Document local management RPC drift as an accepted trade-off. This lets Console converge without waiting for periodic validation. * fix(web): satisfy clippy across managed config sync tests Scope managed runtime guards to blocks in runtime revision tests so no std MutexGuard is held across await points, return the applied revision directly instead of through a let binding, and pass WebhookValidationInput to request_heartbeat_validation instead of expanding it into eight separate arguments. * fix(core): stop reporting failed instances as running in heartbeats A stopped instance with a startup error appeared in both running_network_instances and failed_network_instances, so the server treated it as running and never re-ran its managed config. Exclude failed instance ids when building the running list so the reconciler restarts them. * fix(core): close missed-wakeup race in instance state changes wait_for_change created the Notified future before reading the generation but only registered it when awaited. A change landing in between fired notify_waiters with no registered waiter and delayed the heartbeat by a full interval. Enable the future before reading the generation so every change wakes a waiting heartbeat. * fix(web): address review findings Fence webhook validation and connection transitions against stale state, redact credentials from default-level logs, and stabilize runtime reconciliation: - Record connected bindings only while the session still owns the machine route, and skip disconnect compensation once a replacement owns the route so a stale disconnect cannot revoke it. - Discard webhook validation results when the change epoch moved during the HTTP round, so a stale rejection cannot invalidate the current session. - Drop user_token fields from info and warn logs that became visible with info-level defaults. - Restore a hostname omitted by the 2.6.4 readback into the cached runtime config after a successful mutation, so later rounds stop re-sending the same hostname patch. - Reconcile running web configs when no revision is tracked so legacy unrevisioned updates converge, and wake sessions for unrevisioned full updates instead of waiting for the next heartbeat. * chore(go): regenerate web proto bindings for heartbeat fields Add failed_network_instances, support_heartbeat_policy, and the heartbeat policy response fields to the checked-in Go bindings. Other proto packages are left as-is because their drift predates this change. * fix(web): redact user tokens from positional log arguments Three runtime reconciliation info logs and the user lookup error contexts printed user_token through format arguments, which the earlier field-syntax redaction missed. The reconcile log now fires every round for unrevisioned machines, so remove the token from these messages as well. * fix(web): fence stale validation and runtime reconcile rounds Check webhook validation epochs while holding the session write lock, so stale success and rejection responses cannot change session state. Advance the runtime epoch for unrevisioned full config updates, and exclude failed instances from heartbeat and RPC reconciliation lists so stopped instances are restarted instead of repeatedly hot-patched. Release test read guards before awaiting validation apply calls. Set up the no-pending condition before asserting that an applied revision is a no-op, and verify that its runtime epoch remains unchanged. Validation: all 137 client_manager tests passed. * test(credentials): cover P2P with active VPN portal Model an admin and temporary credential peer connected as a foreign network through a public server with data relay disabled. Verify their direct connection can be replaced after a WireGuard portal client comes online. * test(credentials): stabilize two-admins failover assertions The two-admins non-reusable credential test could fail on slow convergence: after dropping the winning peer it relied on a single route sample passing a bare AND condition, then re-asserted the same expectations through one-shot checks seconds later. A transient route flap in that window (for example a briefly resurrected winner route from stale conn info) turned a passing convergence into a hard assert failure. This matches the 48.9s CI flake of credential_non_reusable_across_two_admins_allows_only_one_peer observed on 2026-08-12. Changes: - wait for bidirectional admin connectivity (AND) with a 20s budget before issuing the credential, instead of a one-directional OR - replace the failover wait_for_condition with wait_stable_failover_visibility_on_admins, which requires three consecutive samples of loser-present and winner-absent on both admins within the same 60s budget and logs every sample - enrich the stable-single-winner timeout message with per-admin visibility flags and elapsed time for triage All existing contracts are preserved; only observation windows and diagnostics change. Validated in the rust container: three passes at normal speed (54.1s / 53.8s / 53.1s) plus one slow-convergence round (172.7s) that would have raced the old one-shot sampling; it now passes with failover samples logged. cargo fmt and clippy -D warnings clean.
This commit is contained in:
1 parent
e0bdb516b6
commit
c96b6c1961
21 files changed
+4332
-561
No files matched your search
@@ -7,9 +7,9 @@
|
||||
- 上游依赖:后续由 Console 计算并发送 Patch
|
||||
- 兼容要求:保留现有 Full PUT
|
||||
|
||||
本文记录当前接收端方案。Session 在能够证明 Patch base 与已应用 revision 连续时
|
||||
只收敛 touched instances;重启、通知丢失、revision 断链或并发积压时沿用 Full
|
||||
reconcile。
|
||||
本文记录当前接收端方案。Session 合并已持久化 Patch 的 touched instance IDs,
|
||||
并在运行态收敛时读取这些实例的最新持久化状态。重启、通知丢失或无法安全判断
|
||||
实例 ownership 时沿用 Full reconcile。
|
||||
|
||||
## 1. 背景与结论
|
||||
|
||||
@@ -30,8 +30,10 @@ Console 每次发布都会向该路径发送完整 Exact Set。实例很多时
|
||||
3. PATCH 使用 `expected_config_revision` 做 compare-and-swap(CAS)。
|
||||
4. Full/Patch 的配置变更与 revision 更新在一个 SQLite transaction 中提交。
|
||||
5. Patch 只查询和写入 touched instances,不扫描完整 Target。
|
||||
6. 写入成功后通知 Session 本次 base、target 和 touched instance IDs。
|
||||
7. Session 仅在 applied revision 精确匹配 base 时增量收敛,否则安全回退 Full。
|
||||
6. 写入成功后通知 Session 本次 expected、target 和 transaction 实际 touched
|
||||
instance IDs。
|
||||
7. Session 只合并 revision 连续的 touched IDs,并以 SQLite 当前状态为准增量
|
||||
收敛;可信 runtime base、通知链或 persisted target 无法证明连续时回退 Full。
|
||||
|
||||
普通变更的接收端成本由:
|
||||
|
||||
@@ -371,20 +373,33 @@ revision。只影响 user-owned rows 的操作不清除 managed revision。
|
||||
- 只有带 target revision 的 `Applied` 才通知匹配的 live Session;
|
||||
`AlreadyApplied`、legacy unrevisioned Full、conflict 和失败不重复通知。
|
||||
- Notification 必须发生在 commit 之后。
|
||||
- Full notification 清除任何 pending delta,触发完整收敛。
|
||||
- Full notification 将 pending reconcile hint 提升为 Full,触发完整收敛。
|
||||
- Patch notification 携带 expected revision、target revision、upsert IDs 和本次
|
||||
transaction 实际接受删除的 web-owned IDs。请求删除但数据库原本不存在的 ID
|
||||
仍是 no-op,不能借机删除 Core 中同 ID 的 user-owned 实例。只有 Session applied
|
||||
revision 精确等于 expected revision,且没有更早的 Patch 等待处理时,才保留该
|
||||
delta。
|
||||
- 两次 Patch 在前一次完成前积压时不合并 delta;Session 清除 pending delta,并在
|
||||
最新 heartbeat/revision 上执行一次 Full。这避免引入 Patch queue 或 delivery FSM。
|
||||
- 增量 round 只读取 upsert rows,只删除本次 delete IDs,只对 touched running
|
||||
instances 执行 runtime Patch/Run。完成前再次校验 persisted target revision;只有
|
||||
全部 touched instances 成功且 target 仍相同,才推进 applied revision。
|
||||
仍是 no-op,不能借机删除 Core 中同 ID 的 user-owned 实例。
|
||||
- Session 将尚未应用的 Patch touched IDs 合并为一个 Dirty set,并始终以 SQLite
|
||||
最新 revision 下的 rows 为准。它不重放历史 Patch,也不维护 Patch queue 或
|
||||
delivery FSM。只有 incoming expected 等于 pending target 的通知才能合并;Dirty
|
||||
hint 保留最早 expected 和最新 target。乱序、不连续或无法证明顺序的通知将 hint
|
||||
提升为 Full。多个连续 Patch 积压时,旧 round 由 runtime epoch 拦截,下一 round
|
||||
直接收敛到最新 target。
|
||||
- Session 分开记录对外报告的 applied revision 和内部可信的 runtime base。开始任何
|
||||
runtime side effect 前清除 applied;Patch round 的 side effects 完全包含在 Dirty
|
||||
set 中,因此失败或被新通知拦截时仍保留最早 runtime base,以便按最新持久化状态
|
||||
重试 Dirty set。Full round、direct mutation、授权失败或 Session ownership 中断会
|
||||
清除 runtime base。
|
||||
- 只有可信 runtime base 等于 Dirty 最早 expected,并且 SQLite persisted revision
|
||||
等于 Dirty 最新 target 时,才允许增量 round。重连后 runtime base 未知、通知
|
||||
丢失,或 SQLite 已经提交了更靠后的 revision 而通知尚未送达时都回退 Full,避免
|
||||
不完整的 Dirty set 把完整 target revision 误标为已应用。
|
||||
- 增量 round 逐个读取 Dirty set 中的最新 row。仍然存在且启用的 web-owned row
|
||||
使用其最新 config;已经删除的 row 进入 delete set;遇到 disabled 或非 web-owned
|
||||
row 时回退 Full,以保留 ownership 规则。完成前再次校验 persisted target
|
||||
revision;只有全部 touched instances 成功且 target 仍相同,才推进 applied
|
||||
revision。
|
||||
- 任何通过 EasyTier Web mutation route 直接 Run、Save、Delete 或切换实例状态的
|
||||
操作在执行前和结束后(包括部分 side effect 后返回错误)都清除 Session applied
|
||||
revision 与 pending delta、增加运行配置 cache epoch,并唤醒一次 Full
|
||||
revision、可信 runtime base 与 pending hint,增加运行配置 cache epoch,并唤醒一次 Full
|
||||
reconcile。旧 round 只有 epoch 仍匹配时才能推进 applied revision;新一轮不得
|
||||
信任 mutation 前缓存的 runtime config。否则 runtime-only mutation 或 Core 成功、
|
||||
SQLite 失败的复合 mutation 可能在 persisted revision 不变时破坏 Patch base 的
|
||||
@@ -491,11 +506,13 @@ response 当作旧 receiver 并静默换一种 mutation contract;出现 404
|
||||
- 超限 Full 稳定返回 413/422,而不是耗尽进程内存;
|
||||
- 并发请求无 deadlock,且 CAS 结果确定。
|
||||
|
||||
Session 测试还必须验证:精确 base/target 使用 touched-instance reconcile;base
|
||||
不匹配、目标 revision 已变化、Full notification 和 Patch backlog 都使用 Full;
|
||||
touched runtime apply 失败不推进 applied revision;删除只作用于本次 delete IDs。
|
||||
运行态 Config Get/Patch/Run/Delete 数量应随 touched instances 增长。为确认运行实例
|
||||
身份而进行的一次 list/meta RPC 可以保留,它不发送或重写所有实例配置。
|
||||
Session 测试还必须验证:连续 Patch 的 Dirty IDs 会合并且保留最早 expected;未知
|
||||
或不匹配的 runtime base、乱序/不连续通知使用 Full;增量 round 读取最新 row;已经
|
||||
删除的 web-owned row 只删除对应 Dirty ID;Full notification 覆盖 Dirty hint;目标
|
||||
revision 已变化或 touched runtime apply 失败时不推进 applied revision;直接 runtime
|
||||
mutation 使 revision 与运行配置 cache 同时失效。运行态 Config Get/Patch/Run/Delete
|
||||
数量应随 touched instances 增长。为确认运行实例身份而进行的一次 list/meta RPC
|
||||
可以保留,它不发送或重写所有实例配置。
|
||||
|
||||
## 11. Observability
|
||||
|
||||
@@ -523,9 +540,13 @@ Rollout acceptance:
|
||||
|
||||
### 12.1 Session runtime delta apply(已实现)
|
||||
|
||||
Patch commit outcome 已携带 touched IDs。Session 只在 applied revision 正好等于
|
||||
Patch base 时执行 touched-instance reconcile;重启、revision 断链、通知丢失或
|
||||
并发 Patch backlog 都退回 Full。接收端不保存 Patch queue,也不合并 delta。
|
||||
Patch commit outcome 已携带 expected、target 和 transaction 实际 touched IDs。
|
||||
Session 只合并 expected/target 连续的 Dirty IDs,并在每一轮从 SQLite 读取最新
|
||||
target revision 对应的当前 rows;因此正常积压只增加 Dirty set,不需要保留中间
|
||||
revision 的 Patch queue。可信 runtime base 必须等于 Dirty 最早 expected;Patch
|
||||
side effect 失败可保留该 base 重试,通知丢失、乱序、进程重启或新 Session 尚无
|
||||
runtime base 时回退 Full。Full
|
||||
notification、disabled row 或 ownership 无法证明时也回退 Full。
|
||||
|
||||
### 12.2 Chunked Full
|
||||
|
||||
|
||||
Reference in new issue
Block a user