mirror of
https://github.com/EasyTier/EasyTier.git
synced 2026-10-08 10:56:13 -08:00
* fix(web): fence managed config runtime reconciliation Keep runtime reconciliation tied to the currently authorized session so stale connections cannot mutate a replacement session runtime. Accumulate only contiguous dirty IDs and load their latest SQLite state. Require the applied revision to match the earliest Patch base and the persisted revision to match the latest target. Otherwise, reconcile the full desired state. Use separate runtime-state and config-cache epochs. Managed updates can reuse observed configs; direct mutations invalidate them. Update sync documentation to match. * fix(web): interrupt validation retry on state changes Track meaningful validation state changes separately from periodic dirty signals. Applied revision changes wake a failed validation immediately, while heartbeat-driven revalidation retains the retry backoff. Treat Notify as a wake-up hint and recheck the state-change epoch after every wake so stored permits and periodic heartbeats cannot cause retry storms. * fix(web): retry unconfirmed connected webhooks Retry node-connected webhook delivery on retryable errors with a short 100ms/500ms backoff and give up immediately on non-retryable errors. Re-check that the session still owns the connection before every attempt and before recording the delivery, so a replaced session can no longer record a stale connected binding. * fix(web): fence disconnects by session ownership Return whether session removal actually removed the current route owner, and emit disconnected only for that owner. Replaced sessions can no longer invalidate a newer connected route. * fix(web): hot-patch managed hostnames Include hostname changes in the hot-patch path instead of falling back to a full restart. When a full overwrite run is required and the desired config has no hostname, inherit the current runtime hostname so an unmanaged value survives until it is explicitly cleared. Read back the runtime config after an overwrite run and verify it converged instead of assuming the desired state was applied. * fix(web): retry transient runtime reconciliation failures Keep the per-session managed runtime reconciliation worker alive when a single database round fails. Retry from the next heartbeat so persisted managed revisions can still converge after restart-time contention. Reserve terminal worker shutdown for destroyed session or storage state, and cover recovery after a transient revision read failure. * fix(web): accept omitted hostname after runtime apply Release 2.6.4 omits hostname from config readback when it matches the device hostname. Trust a successful hostname mutation only when the returned field is absent, while continuing to verify every other field and rejecting explicit mismatches. * fix(web): ignore unmanaged runtime device names Windows release 2.6.4 generates a random interface name when the managed config leaves dev_name empty. Exclude that runtime-owned value from reconciliation unless the desired config explicitly sets a non-empty device name, preventing endless overwrite restarts. * feat(web): report failed network instances to console Expose stopped Core instances with startup errors in heartbeats. Merge Core failures with direct managed-run RPC failures in easytier-web. Send failed instance IDs during token validation without error text. Prune local run failures when managed configs are deleted. * fix(web): distinguish unknown runtime application state Track whether the current session has observed its applied revision separately from the optional revision value. Report this fact through validate-token so Console can preserve application state across receiver restarts while recognizing deliberate pending mutations. * feat(web): configure heartbeat timing from server Heartbeat responses now provide the interval and RPC timeout. Legacy servers use local defaults and remote values are clamped. Web configuration and session receive timeout follow the policy. * fix(web): reject inactive control sessions Route control RPCs by machine id only to sessions whose RPC manager is still running, so a session that has been stopped or replaced can no longer receive control traffic addressed to the device. * fix(core): filter network info before collection When a collect-network-info request names specific instances, collect those instances only instead of collecting every instance and filtering the result afterwards, so unrequested instances no longer run per-collection work on every request. * feat(web): enable focused runtime diagnostics Enable easytier-web info logs by default while preserving explicit log configuration. Record startup settings, session lifecycle, failed instance changes, webhook queue and request latency, and managed runtime operation timings for production diagnosis. * fix(web): preserve managed revision across reconnects Keep one runtime identifier for each Core WebClient lifetime. Reuse its managed runtime state after transport reconnects. Retain applied revisions and reconcile hints while disconnected. Preserve runtime epochs so stale work cannot mark a revision applied. Reject stale sessions from reclaiming routes after reconnect. Core or Web restarts and legacy clients still use unknown state. Immediately revalidate a restored revision after authentication. Document local management RPC drift as an accepted trade-off. This lets Console converge without waiting for periodic validation. * fix(web): satisfy clippy across managed config sync tests Scope managed runtime guards to blocks in runtime revision tests so no std MutexGuard is held across await points, return the applied revision directly instead of through a let binding, and pass WebhookValidationInput to request_heartbeat_validation instead of expanding it into eight separate arguments. * fix(core): stop reporting failed instances as running in heartbeats A stopped instance with a startup error appeared in both running_network_instances and failed_network_instances, so the server treated it as running and never re-ran its managed config. Exclude failed instance ids when building the running list so the reconciler restarts them. * fix(core): close missed-wakeup race in instance state changes wait_for_change created the Notified future before reading the generation but only registered it when awaited. A change landing in between fired notify_waiters with no registered waiter and delayed the heartbeat by a full interval. Enable the future before reading the generation so every change wakes a waiting heartbeat. * fix(web): address review findings Fence webhook validation and connection transitions against stale state, redact credentials from default-level logs, and stabilize runtime reconciliation: - Record connected bindings only while the session still owns the machine route, and skip disconnect compensation once a replacement owns the route so a stale disconnect cannot revoke it. - Discard webhook validation results when the change epoch moved during the HTTP round, so a stale rejection cannot invalidate the current session. - Drop user_token fields from info and warn logs that became visible with info-level defaults. - Restore a hostname omitted by the 2.6.4 readback into the cached runtime config after a successful mutation, so later rounds stop re-sending the same hostname patch. - Reconcile running web configs when no revision is tracked so legacy unrevisioned updates converge, and wake sessions for unrevisioned full updates instead of waiting for the next heartbeat. * chore(go): regenerate web proto bindings for heartbeat fields Add failed_network_instances, support_heartbeat_policy, and the heartbeat policy response fields to the checked-in Go bindings. Other proto packages are left as-is because their drift predates this change. * fix(web): redact user tokens from positional log arguments Three runtime reconciliation info logs and the user lookup error contexts printed user_token through format arguments, which the earlier field-syntax redaction missed. The reconcile log now fires every round for unrevisioned machines, so remove the token from these messages as well. * fix(web): fence stale validation and runtime reconcile rounds Check webhook validation epochs while holding the session write lock, so stale success and rejection responses cannot change session state. Advance the runtime epoch for unrevisioned full config updates, and exclude failed instances from heartbeat and RPC reconciliation lists so stopped instances are restarted instead of repeatedly hot-patched. Release test read guards before awaiting validation apply calls. Set up the no-pending condition before asserting that an applied revision is a no-op, and verify that its runtime epoch remains unchanged. Validation: all 137 client_manager tests passed. * test(credentials): cover P2P with active VPN portal Model an admin and temporary credential peer connected as a foreign network through a public server with data relay disabled. Verify their direct connection can be replaced after a WireGuard portal client comes online. * test(credentials): stabilize two-admins failover assertions The two-admins non-reusable credential test could fail on slow convergence: after dropping the winning peer it relied on a single route sample passing a bare AND condition, then re-asserted the same expectations through one-shot checks seconds later. A transient route flap in that window (for example a briefly resurrected winner route from stale conn info) turned a passing convergence into a hard assert failure. This matches the 48.9s CI flake of credential_non_reusable_across_two_admins_allows_only_one_peer observed on 2026-08-12. Changes: - wait for bidirectional admin connectivity (AND) with a 20s budget before issuing the credential, instead of a one-directional OR - replace the failover wait_for_condition with wait_stable_failover_visibility_on_admins, which requires three consecutive samples of loser-present and winner-absent on both admins within the same 60s budget and logs every sample - enrich the stable-single-winner timeout message with per-admin visibility flags and elapsed time for triage All existing contracts are preserved; only observation windows and diagnostics change. Validated in the rust container: three passes at normal speed (54.1s / 53.8s / 53.1s) plus one slow-convergence round (172.7s) that would have raced the old one-shot sampling; it now passes with failover samples logged. cargo fmt and clippy -D warnings clean.
79 lines
5.2 KiB
YAML
79 lines
5.2 KiB
YAML
_version: 2
|
|
|
|
cli:
|
|
db:
|
|
en: "path to the sqlite3 database file, used to save all the data"
|
|
zh-CN: "sqlite3 数据库文件路径, 用于保存所有数据"
|
|
console_log_level:
|
|
en: "The log level for the console logger. Possible values: trace, debug, info, warn, error"
|
|
zh-CN: "控制台日志级别。可能的值:trace, debug, info, warn, error"
|
|
file_log_level:
|
|
en: "The log level for the file logger. Possible values: trace, debug, info, warn, error"
|
|
zh-CN: "文件日志级别。可能的值:trace, debug, info, warn, error"
|
|
file_log_dir:
|
|
en: "The directory to save the log files, default is the current directory"
|
|
zh-CN: "保存日志文件的目录,默认为当前目录"
|
|
config_server_port:
|
|
en: "The port to listen for the config server, used by the easytier-core to connect to"
|
|
zh-CN: "配置服务器的监听端口,用于被 easytier-core 连接"
|
|
config_server_protocol:
|
|
en: "The protocol to listen for the config server, used by the easytier-core to connect to, possible values: udp, tcp, ws"
|
|
zh-CN: "配置服务器的监听协议,用于被 easytier-core 连接, 可能的值:udp, tcp, ws"
|
|
api_server_port:
|
|
en: "The port to listen for the restful server, acting as ApiHost and used by the web frontend"
|
|
zh-CN: "restful 服务器的监听端口,作为 ApiHost 并被 web 前端使用"
|
|
api_server_addr:
|
|
en: "The listen address for the restful server, e.g. 0.0.0.0, ::, 127.0.0.1"
|
|
zh-CN: "restful 服务器的监听地址, 例如 0.0.0.0, ::, 127.0.0.1"
|
|
web_server_port:
|
|
en: "The port to listen for the web dashboard server, default is same as the api server port"
|
|
zh-CN: "web dashboard 服务器的监听端口, 默认为与 api 服务器端口相同"
|
|
web_server_addr:
|
|
en: "The listen address for the web dashboard server (only effective when web_server_port differs from api_server_port or web_server_addr differs from api_server_addr), e.g. 0.0.0.0, ::, 127.0.0.1"
|
|
zh-CN: "web dashboard 服务器的监听地址(仅在 web_server_port 与 api_server_port 不同,或 web_server_addr 与 api_server_addr 不同时生效), 例如 0.0.0.0, ::, 127.0.0.1"
|
|
no_web:
|
|
en: "Do not run the web dashboard server"
|
|
zh-CN: "不运行 web dashboard 服务器"
|
|
api_host:
|
|
en: "The URL of the API server, used by the web frontend to connect to"
|
|
zh-CN: "API 服务器的 URL,用于 web 前端连接"
|
|
geoip_db:
|
|
en: "The path to the GeoIP2 database file, used to lookup the location of the client, default is the embedded file (only country information) , recommend https://github.com/P3TERX/GeoLite.mmdb"
|
|
zh-CN: "GeoIP2 数据库文件路径,用于查找客户端的位置,默认为嵌入文件(仅国家信息),推荐 https://github.com/P3TERX/GeoLite.mmdb"
|
|
heartbeat_min_response_ms:
|
|
en: "Config-server heartbeat interval in milliseconds, default is 3500"
|
|
zh-CN: "配置服务心跳周期,单位毫秒,默认为 3500"
|
|
heartbeat_timeout_ms:
|
|
en: "Config-server heartbeat RPC timeout in milliseconds, default is 15000"
|
|
zh-CN: "配置服务心跳 RPC 超时时间,单位毫秒,默认为 15000"
|
|
disable_registration:
|
|
en: "Disable user registration"
|
|
zh-CN: "禁用用户注册"
|
|
oidc_issuer_url:
|
|
en: "The OIDC issuer URL for single sign-on authentication"
|
|
zh-CN: "OIDC 签发者 URL,用于单点登录认证"
|
|
oidc_client_id:
|
|
en: "The OIDC client ID"
|
|
zh-CN: "OIDC 客户端 ID"
|
|
oidc_client_secret:
|
|
en: "The OIDC client secret (can also be set via OIDC_CLIENT_SECRET env var)"
|
|
zh-CN: "OIDC 客户端密钥(也可通过 OIDC_CLIENT_SECRET 环境变量设置)"
|
|
oidc_username_claim:
|
|
en: "The OIDC claim to use as the local username, default: preferred_username"
|
|
zh-CN: "用作本地用户名的 OIDC claim 字段,默认: preferred_username"
|
|
oidc_scopes:
|
|
en: "OIDC scopes to request during login. Supports comma-separated values or repeated --oidc-scopes flags, default: openid,profile"
|
|
zh-CN: "登录时请求的 OIDC scopes。支持逗号分隔或多次指定 --oidc-scopes,默认: openid,profile"
|
|
oidc_redirect_url:
|
|
en: "The OIDC redirect URL (callback URL), must match exactly what is registered with your Identity Provider. Required when using OIDC. Example: http://your-domain.com:11211/api/v1/auth/oidc/callback"
|
|
zh-CN: "OIDC 重定向 URL(回调 URL),必须与身份提供商注册的地址完全一致。使用 OIDC 时必须提供。示例: http://your-domain.com:11211/api/v1/auth/oidc/callback"
|
|
allow_auto_create_user:
|
|
en: "Allow auto-creating local user when easytier-core connects with an unknown username"
|
|
zh-CN: "当 easytier-core 使用未知用户名连接时,允许自动创建本地用户"
|
|
oidc_disable_pkce:
|
|
en: "Disable PKCE (Proof Key for Code Exchange) for OIDC authentication"
|
|
zh-CN: "禁用 OIDC 认证的 PKCE(授权码交换证明密钥)"
|
|
oidc_frontend_base_url:
|
|
en: "Frontend base URL to redirect to after successful OIDC callback. Required when frontend and API are deployed separately (non-embed build, --no-web mode, or different web_server_port)"
|
|
zh-CN: "OIDC 回调成功后跳转的前端入口地址。当前端与 API 分离部署时必须提供(非 embed 构建、--no-web 模式、或 web_server_port 与 api_server_port 不同)"
|