Commit Graph
16 Commits
Author SHA1 Message Date
KKRainbow 651a8d9e25 feat(web): add central network management console (#2622)
Persist central network intent and compile complete per-device configs
with secure credentials, ACL policy, and dedicated Gateway runtimes.
Expose network, credential, device registry, and runtime observation
APIs.

Treat central management and external Console as alternative consumers
of the upstream ClientManager. Register central devices through a local
webhook handler and publish through existing runtime reconciliation.
Serialize public mutations with enrollment, reject direct credential
changes to managed instances, and retain REST revision invalidation.

Keep Gateway lifecycle publication in the central service and remove
obsolete incremental result bookkeeping. Restore persisted networks and
retire orphan runtimes through the same serialized publisher.

Add Core and protocol support for Gateway and WireGuard management,
including GUI bindings and serialized GUI config writes. Bootstrap IPv4
for DHCP-only networks and retain assigned addresses without peer IPv4.

Cover intent transactions, complete configuration publication, offline
recovery, authentication, revocation, Gateway lifecycle, and DHCP.

Validate central candidates before persistence using Core URL, config,
and portal-client rules. Reject unsupported peer schemes, unconvertible
proxy subnets, and invalid portal clients without changing live intent.
Expose existing pure Core validators without changing runtime behavior.

Gate API-facing credential validation on API-enabled Core builds so
minimal WASM targets do not reference omitted management types.

Preserve session-backed device views in external Console mode. Bound
Gateway admissions without cancelling transport upgrades, and cancel
pending peer handshakes before retiring runtimes. Batch registry reads
and bound runtime observation concurrency.

Randomize DHCP bootstrap and conflict retries, keep an advertised
current subnet, and select fallback subnets deterministically. Cover
concurrent startup, Gateway admission lifetime, and external Console
regressions.

Allow DHCP bootstrap with no remote routes. Count network members
through one tenant-scoped query without write transactions. Cover zero-
peer allocation and tenant/empty-network counts.

Preserve direct Web configurations and disabled states in central device
snapshots. Remove obsolete central rows atomically with membership
changes. Stop online managed instances before deleting or blocking
devices, retaining intent when shutdown or deletion fails.

Persist automatic IPv4 allocations separately from manual overrides so
subnet changes reallocate only automatic members. Normalize mapped proxy
routes to their advertised CIDRs when granting temporary credentials.
Cover publication, deletion rollback, subnet migration, and grant
updates.

Read central intent in deferred transactions so polling does not reserve
the SQLite writer lock. Test reads with an active writer and assert the
temporary-member secret constraint using a valid device fixture.

Persist patched WireGuard clients from the saved Web or GUI candidate
without a follow-up RPC. Preserve pending settings and ownership, and
cover disconnects and failed patches.

Resolve named credential and config mutations against the live Core
instance before checking central ownership. Forward authorized mutations
by UUID on the same session, including while network renames are
pending. Cover all mutation methods and preserve direct Console
behavior.

* feat(web-ui): add central network console and WireGuard management

Add network and device views with central membership, credentials, ACL
policy, temporary peers, and per-instance runtime details. Update
console navigation, styling, theme handling, and API clients.

Extend shared configuration and status components for central networks.
Add a WireGuard portal dialog for setup and running-device management,
with matching translations and network configuration types.

Use secure UUID generation on HTTP, discard stale member configuration
responses, and stop node-detail polling when the component unmounts.

Update frontend workspace dependencies and include component, dashboard,
configuration serialization, and central console end-to-end tests.

Add isolated real-Core E2E coverage for central data-plane traffic, ACL,
DHCP, recovery, device lifecycle, and native WireGuard clients. Record
all 59 functional checks with evidence and confirmed validation/UI
defects; keep runtime artifacts and test credentials out of Git.

Normalize protobuf logger levels and render translated labels correctly.
Cover all six levels across setting, reload, and language changes. Add
real configuration-rejection E2E checks for database and Core stability
and uninterrupted traffic, and record the resolved audit findings.

Gate central navigation and registry actions by console mode. Refresh
WireGuard settings on mounted status views and preserve explicit portal
listener endpoints. Explain the trusted permanent-member ACL boundary.

Cover external Console rendering, portal refresh retries and teardown,
and explicit IPv4/IPv6 listener exports.

Preserve PublicServer discovery when saving its settings, including when
its URL matches the Gateway. Cover renaming and endpoint edits in the
browser.

Display automatic member addresses without converting them to manual
overrides during edits. Derive offline address prefixes from the network
subnet. Add browser regression coverage and real-Core checks for direct
configuration preservation, temporary proxy mappings, automatic subnet
migration, and shutdown before device deletion.

Build enrollment commands from the configured API hostname, including
IPv6 and relative endpoints. Use the PublicServer connector in temporary
credential CLI and TOML exports. Add browser regression coverage.

Preserve form credentials across repeated normalization and GUI storage
reloads. Keep explicit form values authoritative over backend keys and
cover JSON persistence, idempotence, replacement, and clearing.
2026-10-05 11:06:47 +08:00
KKRainbow 3da349f4ee fix(relay): preserve forwarded packets in secure relay paths (#2618)
Forward remote packets without starting a relay session at the
intermediate hop. Encrypting forwarded Noise handshake replies with
that hop's session corrupts the end-to-end handshake and causes
credential connectivity tests to time out.

Disable automatic P2P in the admin/shared/credential test so direct
connections cannot mask broken forwarding. Cover unchanged forwarding
of handshake requests, replies, and encrypted data packets.
2026-09-29 22:52:10 +08:00
KKRainbow ff3921ce68 docs: acknowledge Linfeng Cloud in sponsor sections (#2604)
Add Linfeng Cloud to both README sponsor sections, using the
partner link from the project website. Store its logo locally so it
renders on GitHub without relying on the partner site.
2026-09-24 12:05:36 +08:00
KKRainbow ebf0b947b9 fix(ospf): resync missing peer data after metadata expiry (#2599)
Report missing peer metadata and connection rows in sync responses.
Resend those records through the existing per-session pending queues
without resetting the incremental cursor or rebuilding sessions.

Keep requests across lost responses, distinguish missing connection rows
from metadata, and avoid requesting non-relaying credential rows.
Deduplicate records selected by both incremental and pending paths.

Cover selective expiry, multihop recovery, retries, partial availability,
and unchanged-record suppression. All 40 route module tests pass.

Fixes #2597
2026-09-22 23:38:34 +08:00
KKRainbow 0f3d8e4434 fix(core): keep peer liveness responsive under receive limits (#2590) 2026-09-19 08:22:12 +08:00
KKRainbow e313ba8efb fix(wasi): reduce startup RSS, align ABI v4 and wire v3 (#2588)
* feat(easytier-go): accept data plane ABI v4

The core raised DATA_PLANE_ABI_VERSION to 4 in f26c2aa1 ("feat(wasi):
run EasyTier core on Cloudflare Workers and browsers") to advertise the
new guest exports easytier_data_plane_tcp_shutdown_write_submit and
easytier_data_plane_tcp_shutdown_write_result_take, which let a host
half-close guest TCP streams.

The Go host has no caller for half-close: net.Conn exposes only Close,
which already tears down both directions, so the v3 behavior is
preserved and no new plumbing is added. Without this bump, any artifact
rebuilt from current core source is rejected at instance creation with
"unsupported EasyTier data plane ABI version 4, want 3".

* fix(easytier-go): decode socket options wire v3, rebuild core artifact

38e2a621 ("refactor(ohos): 拆分 OHRS 包并按 socket 精细保护 VPN 流量")
raised the host socket options wire format from version 2 to 3: a
need_protect byte is appended after the purpose byte in TCP connect,
UDP bind, and TCP listen options, shifting bind_device one byte later.
The Go hostABI decoders were never updated, so every socket operation
from a HEAD-built core was rejected as "invalid options" and required
listeners failed to start.

Bump the accepted wire version to 3 and skip the need_protect byte in
all three decoders. The byte requests VPN socket protection, which only
Android-style VPN hosts can honor; on every other platform sockets are
already protected, so reading and discarding it is correct.

Regenerate the embedded core artifact and protobuf bindings from HEAD
(599e4eac) so the shipped artifact matches the Go host again. The proto
regeneration also picks up schema fields added since the last embed
(e.g. prefer_peer_relay).

* perf(easytier-go): release compiler garbage after host init

wazero's optimizing compiler allocates ~100MB of throwaway state on
the Go heap while compiling the embedded core. Go's runtime does not
return that memory to the OS after the initiating GC, so the process
retained the compilation peak for its entire lifetime: RSS sat at
~164MB before any instance or network activity.

Call debug.FreeOSMemory() once after the module is instantiated.
NewHost is a one-time initialization path outside the dataplane, so
the stop-the-world pass is safe here. Measured RSS after host.New
drops from ~164MB to ~64MB; no behavioral change.

* chore(easytier-js): bump toolchain dependencies

- vitest 2.1.9 -> 3.2.7 (all three packages)
- esbuild 0.25.9 -> 0.28.2 (browser bundler)
- wrangler 4.114.0 -> 4.134.0 (cloudflare + web example)
- @cloudflare/workers-types 5.20260724.1 -> 5.20260917.1
- binaryen 131.0.0 -> 132.0.0 (JS bindings only; the wasm build
  uses the standalone wasm-opt binary fetched by the build script)

Not bumped: typescript stays at 5.9.3 (latest 5.x; 7.0 is a major
jump not worth taking for this workspace), vite stays at 5.4.21
(web-example only; 5->8 spans three majors).

Verified with pnpm test (28 + 2 + 2 tests across runtime, browser,
cloudflare) and pnpm check (tsc + wrangler deploy --dry-run).
pnpm-workspace.yaml gained three minimumReleaseAgeExclude entries
recorded automatically by pnpm for freshly published versions.

* fix(easytier-js): align host wire format with core wire v3

The core socket options wire format moved to version 3 in 38e2a621
(need_protect byte after purpose). The JS websocket-host still
required version 2 and rejected every TCP bind from a HEAD-built core,
breaking browser port leases.

- websocket-host.ts: accept version 3, minimum length 49. The
  need_protect byte sits after purpose (offset 43) and is a no-op
  outside Android VPN hosts, so the decoder just skips it.
- websocket-host.test.ts: update the test encoder to emit v3.
- binaryen stays at 131.0.0 to match script/build-wasi-core.sh, which
  intentionally pins binaryen 131 for the Go-side embedded core.
  The npm binaryen package provides the wasm-opt binary used by
  build-wasm.mjs, so keeping both sides on the same version avoids
  divergent optimization output.
- pnpm-workspace.yaml: drop two stale minimumReleaseAgeExclude entries
  for @cloudflare/workers-types versions no longer in the lockfile.

* test(easytier-go): cover socket options wire v3 decoders

Direct unit tests for decodeTCPConnectOptions, decodeUDPBindOptions, and
decodeTCPListenOptions with wire v3 documents. Covers combinations of
socket mark, netns, bind device, and local address presence, plus
rejection of wire v2.

Resolves review feedback on PR #2588.
2026-09-19 00:13:47 +08:00
KKRainbow 7853ec8685 fix(windows): reliable service auto-start on boot and start-on-boot GUI entry (#2579)
* fix(service): restart windows service indefinitely on failure

The windows service was installed with only the AutoStart start type
and no SCM failure actions configured. When the service failed during
boot (e.g. network not yet ready, config load error), it reported
SERVICE_STOPPED with a non-zero exit code and stayed stopped until
started manually, which is reported as no auto-start after reboot in
issue #1771.

Configure failure actions on install and update when restart is not
disabled:

- three Restart actions (1s/5s/10s delay) with reset period Never;
  the SCM repeats the last action once the failure count exceeds the
  actions array, so restarts are retried indefinitely
- set fFailureActionsOnNonCrashFailures so exits that report
  SERVICE_STOPPED with a non-zero exit code (the error path in
  win_service_event_loop) are also treated as failures; manual stops
  still exit with 0 and do not trigger a restart

Also exit the service process with a non-zero code after reporting the
error status. Without it the process stayed alive after reporting
SERVICE_STOPPED and was only counted as failed after the SCM
force-killed it, adding roughly 30s of dead time to every retry cycle.

This matches the systemd path, which already generates Restart=always
with StartLimitIntervalSec=0. The --disable-restart-on-failure option
now also clears previously configured failure actions on windows.

Verified on a real windows host: crash failures and reported-error
failures both restart with the configured 1s/5s/10s cadence
indefinitely, manual stops are not restarted, and
--disable-restart-on-failure clears the actions.

* feat(gui): add start-on-boot menu entry pointing to service mode

Issue #1771 reports that users cannot find how to make EasyTier start
on boot. Auto-start is provided by service mode, but the GUI offered
no entry named after it, so the connection was hard to discover.

Add a "Start on Boot" item to the settings menu. It opens the mode
dialog with service mode preselected and shows an info message
explaining that enabling service mode registers EasyTier as a system
service that starts automatically at boot and keeps running in the
background.

When the dialog is opened with service mode preselected, the mode
watcher in ModeSwitcher can run before the default config/log dirs
have been resolved, leaving the fields empty and failing validation on
save. Fill them from the resolved defaults after mount in that case.

Add mode.autostart / mode.autostart_hint strings to the cn/en locales
in frontend-lib.

* fix(cli): stop collecting the --core-args flag into the service args

InstallArgs.core_args was declared without an explicit `long`, so
clap treated it as a trailing positional argument instead of a named
option. Passing `service install --core-args --daemon ...` therefore
collected the literal "--core-args" token into the value, and the
installed service was registered with an invalid command line that
failed on every start (observed on a real windows host: the binPath
contained `easytier-core.exe --core-args --daemon ...`).

Declare it as a real option (`long` + `num_args = 1..`) while keeping
allow_hyphen_values and the trailing semantics: --core-args must be
the last option of install and consumes everything after it.

The bare-positional spelling (`service install --daemon`), the only
correctly-working form before, now fails with a clear "unexpected
argument" error; scripts written against it need to add the
--core-args prefix.

Add unit tests covering the flag, `=` and mixed forms.
2026-09-16 23:02:40 +08:00
KKRainbow 5477d4bca2 fix(proxy): isolate TCP flows and recover KCP control loss (#2569)
* fix(proxy): distinguish TCP flows sharing source ports

Key TCP NAT entries by source and destination so connections sharing a
source port can reach different targets independently. Allocate a unique
translated source port for each flow and restore the original addresses
and ports on the return path.

Keep connection identity during accept and cleanup so a replaced flow
cannot lose its new mapping. Give each smoltcp proxy mode its own
listener port and retain the original destination for mapped subnets.

Include regression coverage for SYN retries, concurrent insertion,
stale cleanup, translated port wraparound and exhaustion, plus the
recorded cross-platform and six-mode traffic validation.

* fix(proxy): preserve half-closed streams and pending KCP accepts

Propagate EOF to the opposite writer while keeping reverse traffic
alive until both directions finish. This allows a server to reply after
reading request EOF and a client to send after reading response EOF.

Pin kcp-sys to b37ee660 for retained pending accepts, receive-buffer
draining before EOF, early FIN handling, atomic state cleanup, and
keepalives that start only after the handshake completes.

Cover both shutdown directions, empty requests and buffered transfers
in gateway tests and the TCP/KCP/QUIC kernel/smoltcp integration matrix.
Include the half-close validation record and its remaining loss limits.

* fix(proxy): enable backward-compatible KCP control recovery

Pin kcp-sys to 26853356 and update both the root and OHOS workspace
lockfiles. The dependency negotiates reliable control per connection
while preserving legacy behavior when either endpoint lacks support.

New peers recover lost handshake and FIN control packets, acknowledge
FIN without closing the reverse writer, and retain bounded close state
for duplicate packets. Queue saturation cannot block reliable stateless
responses. Compatibility tests retain one historical baseline.

Include the protocol design and validation records for deterministic
loss, real old endpoints, three platforms, the full integration matrix,
and proxy traffic. Preserve mixed-version failures and the documented
minor scheduling issue rather than claiming legacy loss recovery.

 Validate the OHOS workspace with --locked: 12 library tests and the
binding test build check pass, along with formatting.
2026-09-16 12:50:08 +08:00
KKRainbow dd013e6a2a fix(acl): prevent bidirectional ping from bypassing inbound drop rules (#2570)
* fix(acl): match ICMP replies in outbound allow records

Only outbound echo requests with code zero create ICMP response records.
Require inbound complete packets and first fragments to be echo replies
with code zero before using those records, so reverse echo requests
remain subject to inbound ACL rules.

Recognize ICMP fragment tails without reading their payload as an ICMP
header. Allow them to use existing address and protocol records without
creating new records. Walk IPv6 extension headers to locate ICMPv6 and
avoid inferring an unknown tail protocol from its payload.

Add IPv4 and IPv6 unit coverage for message direction, code, truncation,
and fragment parsing. Add a three-node regression for bidirectional
requests and ordinary and fragmented replies under inbound default-drop.

Fixes #2545

* fix(acl): keep IPv6 parse failures subject to rules

Retain source and destination addresses when an IPv6 extension header
is truncated, exceeds the payload, or repeats a fragment header. Treat
these packets as unspecified protocol instead of entering the global
parse-failure allow path. Apply the same handling to truncated transport
headers after an extension header.

Classify non-ICMPv6 fragment tails as unspecified so their data cannot
be parsed as TCP/UDP ports or create temporary response records. Keep
ICMPv6 tail authorization through existing response records.

Exercise the ACL entry point with malformed headers and short and long
fragment tails under default-drop and default-allow policies, including
port allow rules that must not match tail payload bytes.
2026-09-15 10:03:15 +08:00
KKRainbow c96b6c1961 fix(web): harden managed config sync between console and clients (#2567)
* fix(web): fence managed config runtime reconciliation

Keep runtime reconciliation tied to the currently authorized session so
stale connections cannot mutate a replacement session runtime.

Accumulate only contiguous dirty IDs and load their latest SQLite state.
Require the applied revision to match the earliest Patch base and the
persisted revision to match the latest target. Otherwise, reconcile the
full desired state.

Use separate runtime-state and config-cache epochs. Managed updates can
reuse observed configs; direct mutations invalidate them. Update sync
documentation to match.

* fix(web): interrupt validation retry on state changes

Track meaningful validation state changes separately from periodic dirty signals. Applied revision changes wake a failed validation immediately, while heartbeat-driven revalidation retains the retry backoff.

Treat Notify as a wake-up hint and recheck the state-change epoch after every wake so stored permits and periodic heartbeats cannot cause retry storms.

* fix(web): retry unconfirmed connected webhooks

Retry node-connected webhook delivery on retryable errors with a
short 100ms/500ms backoff and give up immediately on non-retryable
errors. Re-check that the session still owns the connection before
every attempt and before recording the delivery, so a replaced
session can no longer record a stale connected binding.

* fix(web): fence disconnects by session ownership

Return whether session removal actually removed the current route owner, and emit disconnected only for that owner. Replaced sessions can no longer invalidate a newer connected route.

* fix(web): hot-patch managed hostnames

Include hostname changes in the hot-patch path instead of falling
back to a full restart. When a full overwrite run is required and
the desired config has no hostname, inherit the current runtime
hostname so an unmanaged value survives until it is explicitly
cleared.

Read back the runtime config after an overwrite run and verify it
converged instead of assuming the desired state was applied.

* fix(web): retry transient runtime reconciliation failures

Keep the per-session managed runtime reconciliation worker alive when a
single database round fails. Retry from the next heartbeat so persisted
managed revisions can still converge after restart-time contention.

Reserve terminal worker shutdown for destroyed session or storage state,
and cover recovery after a transient revision read failure.

* fix(web): accept omitted hostname after runtime apply

Release 2.6.4 omits hostname from config readback when it matches the device hostname. Trust a successful hostname mutation only when the returned field is absent, while continuing to verify every other field and rejecting explicit mismatches.

* fix(web): ignore unmanaged runtime device names

Windows release 2.6.4 generates a random interface name when the managed config leaves dev_name empty. Exclude that runtime-owned value from reconciliation unless the desired config explicitly sets a non-empty device name, preventing endless overwrite restarts.

* feat(web): report failed network instances to console

Expose stopped Core instances with startup errors in heartbeats.

Merge Core failures with direct managed-run RPC failures in easytier-web.

Send failed instance IDs during token validation without error text.

Prune local run failures when managed configs are deleted.

* fix(web): distinguish unknown runtime application state

Track whether the current session has observed its applied revision
separately from the optional revision value. Report this fact through
validate-token so Console can preserve application state across
receiver restarts while recognizing deliberate pending mutations.

* feat(web): configure heartbeat timing from server

Heartbeat responses now provide the interval and RPC timeout.

Legacy servers use local defaults and remote values are clamped.

Web configuration and session receive timeout follow the policy.

* fix(web): reject inactive control sessions

Route control RPCs by machine id only to sessions whose RPC manager
is still running, so a session that has been stopped or replaced
can no longer receive control traffic addressed to the device.

* fix(core): filter network info before collection

When a collect-network-info request names specific instances,
collect those instances only instead of collecting every instance
and filtering the result afterwards, so unrequested instances no
longer run per-collection work on every request.

* feat(web): enable focused runtime diagnostics

Enable easytier-web info logs by default while preserving explicit log configuration. Record startup settings, session lifecycle, failed instance changes, webhook queue and request latency, and managed runtime operation timings for production diagnosis.

* fix(web): preserve managed revision across reconnects

Keep one runtime identifier for each Core WebClient lifetime.

Reuse its managed runtime state after transport reconnects.

Retain applied revisions and reconcile hints while disconnected.

Preserve runtime epochs so stale work cannot mark a revision applied.

Reject stale sessions from reclaiming routes after reconnect.

Core or Web restarts and legacy clients still use unknown state.

Immediately revalidate a restored revision after authentication.

Document local management RPC drift as an accepted trade-off.

This lets Console converge without waiting for periodic validation.

* fix(web): satisfy clippy across managed config sync tests

Scope managed runtime guards to blocks in runtime revision tests so
no std MutexGuard is held across await points, return the applied
revision directly instead of through a let binding, and pass
WebhookValidationInput to request_heartbeat_validation instead of
expanding it into eight separate arguments.

* fix(core): stop reporting failed instances as running in heartbeats

A stopped instance with a startup error appeared in both
running_network_instances and failed_network_instances, so the
server treated it as running and never re-ran its managed config.
Exclude failed instance ids when building the running list so the
reconciler restarts them.

* fix(core): close missed-wakeup race in instance state changes

wait_for_change created the Notified future before reading the
generation but only registered it when awaited. A change landing in
between fired notify_waiters with no registered waiter and delayed
the heartbeat by a full interval. Enable the future before reading
the generation so every change wakes a waiting heartbeat.

* fix(web): address review findings

Fence webhook validation and connection transitions against stale
state, redact credentials from default-level logs, and stabilize
runtime reconciliation:

- Record connected bindings only while the session still owns the
  machine route, and skip disconnect compensation once a replacement
  owns the route so a stale disconnect cannot revoke it.
- Discard webhook validation results when the change epoch moved
  during the HTTP round, so a stale rejection cannot invalidate the
  current session.
- Drop user_token fields from info and warn logs that became
  visible with info-level defaults.
- Restore a hostname omitted by the 2.6.4 readback into the cached
  runtime config after a successful mutation, so later rounds stop
  re-sending the same hostname patch.
- Reconcile running web configs when no revision is tracked so
  legacy unrevisioned updates converge, and wake sessions for
  unrevisioned full updates instead of waiting for the next
  heartbeat.

* chore(go): regenerate web proto bindings for heartbeat fields

Add failed_network_instances, support_heartbeat_policy, and the
heartbeat policy response fields to the checked-in Go bindings.
Other proto packages are left as-is because their drift predates
this change.

* fix(web): redact user tokens from positional log arguments

Three runtime reconciliation info logs and the user lookup error
contexts printed user_token through format arguments, which the
earlier field-syntax redaction missed. The reconcile log now fires
every round for unrevisioned machines, so remove the token from
these messages as well.

* fix(web): fence stale validation and runtime reconcile rounds

Check webhook validation epochs while holding the session write lock,
so stale success and rejection responses cannot change session state.
Advance the runtime epoch for unrevisioned full config updates, and
exclude failed instances from heartbeat and RPC reconciliation lists
so stopped instances are restarted instead of repeatedly hot-patched.

Release test read guards before awaiting validation apply calls. Set
up the no-pending condition before asserting that an applied revision
is a no-op, and verify that its runtime epoch remains unchanged.

Validation: all 137 client_manager tests passed.

* test(credentials): cover P2P with active VPN portal

Model an admin and temporary credential peer connected as a foreign network through a public server with data relay disabled. Verify their direct connection can be replaced after a WireGuard portal client comes online.

* test(credentials): stabilize two-admins failover assertions

The two-admins non-reusable credential test could fail on slow
convergence: after dropping the winning peer it relied on a single
route sample passing a bare AND condition, then re-asserted the same
expectations through one-shot checks seconds later. A transient route
flap in that window (for example a briefly resurrected winner route
from stale conn info) turned a passing convergence into a hard assert
failure. This matches the 48.9s CI flake of
credential_non_reusable_across_two_admins_allows_only_one_peer
observed on 2026-08-12.

Changes:

- wait for bidirectional admin connectivity (AND) with a 20s budget
  before issuing the credential, instead of a one-directional OR
- replace the failover wait_for_condition with
  wait_stable_failover_visibility_on_admins, which requires three
  consecutive samples of loser-present and winner-absent on both
  admins within the same 60s budget and logs every sample
- enrich the stable-single-winner timeout message with per-admin
  visibility flags and elapsed time for triage

All existing contracts are preserved; only observation windows and
diagnostics change. Validated in the rust container: three passes at
normal speed (54.1s / 53.8s / 53.1s) plus one slow-convergence round
(172.7s) that would have raced the old one-shot sampling; it now
passes with failover samples logged. cargo fmt and clippy -D warnings
clean.
2026-09-13 01:13:28 +08:00
KKRainbow e0bdb516b6 fix(quic): bind proxy packet checksum to packet number via ETQ1 version (#2565)
* fix(quic): bind proxy packet checksum to packet number via ETQ1 version

QUIC proxy connections die with quinn PROTOCOL_VIOLATION "unsent
packet acked" under bursty traffic with reordering, and the affected
peer pair keeps failing for every new connection until the source node
restarts.

Root cause: the custom crypto checksums the packet bytes but not the
packet number, while quinn decodes truncated packet numbers by
proximity to the largest received (RFC 9000 Appendix A). A 1-byte
encoded packet delayed beyond the +/-128 decode window is decoded as
a future packet number, still passes the checksum, gets ACKed, and
the peer aborts because it never sent that number. Real QUIC survives
this because the AEAD nonce is derived from the packet number, so a
misdecode fails authentication.

Fix: negotiate a custom QUIC version ETQ1 (0x45545131) for the proxy.
Connections on ETQ1 mix the packet number into the SeaHash checksum,
so an out-of-window misdecode fails authentication and the packet is
handled as ordinary loss. The mode is derived statelessly from the
negotiated version in QuicSession and ServerConfig::initial_keys.

Compatibility: the proxy endpoint accepts both ETQ1 and version 1.
NatDstQuicConnector dials ETQ1 first; on ConnectionError::VersionMismatch
from a legacy peer it retries with version 1 and remembers the peer in
legacy_version_peers to skip the rejected version afterwards. The
quic:// tunnel keeps version 1 only.

Verified with 13 docker nodes under netem jitter and bursty iperf
load: the previously-poisoned pair survived 16 minutes on ETQ1 with
zero violations while all legacy-version pairs kept dying; mixed
new-to-old and old-to-new connections work.

* fix(quic): extend the ETQ1 checksum fix to the quic:// tunnel

The quic:// tunnel shares CryptoKey with the proxy, so after the proxy
moved to ETQ1 the tunnel still carried the unsent-packet-acked exposure.
Make endpoint_config() dual-version so tunnel listeners accept both
ETQ1 and legacy peers, and dial ETQ1 first in upgrade_connected with a
transparent fallback to version 1 on VersionMismatch, via a shared
connect_with_etq1 helper. The proxy keeps its hedged dialer with the
per-peer legacy memory; tunnel connections are established once per
session, so the fallback there costs a single extra round trip.
2026-09-11 23:33:46 +08:00
KKRainbow 19e5c49ba3 chore: bump version to 2.7.0 (#2566)
* chore: bump version to 2.7.0

Update version strings across the workspace:
- crate versions and internal dependency requirements for
  easytier, easytier-core, easytier-proto, easytier-web,
  easytier-gui, and easytier-mini, plus Cargo.lock
- GUI package.json and tauri.conf.json
- Magisk module.prop
- default release/docker workflow tags (v2.7.0)

* fix(ohos): sync easytier-ohrs Cargo.lock with bumped workspace versions

The ohos workflow builds easytier-ohrs with --locked, so its
lockfile must record the new 2.7.0 versions of the easytier,
easytier-core, and easytier-proto path dependencies.
2026-09-10 22:38:24 +08:00
KKRainbow 974c3270e8 feat(magisk): add module WebUI configuration (#2563)
* feat(magisk): add module WebUI configuration

Reuse the existing config generator in KernelSU-compatible module
managers. Validate and atomically persist TOML through a module helper
before restarting EasyTier, and package the generated assets in CI.

Closes #1915

* fix(magisk): preserve existing WebUI configuration

Merge form-managed fields into the original TOML so advanced module settings remain intact. Resolve the running core by its exact executable path before restart.
2026-09-10 20:13:21 +08:00
KKRainbow 3d0c9c3ca5 chore(go): use lowercase module import path (#2560)
* chore(go): use lowercase module import path
* ci: scope checks for Go module changes
2026-09-10 15:13:55 +08:00
KKRainbow 86d942ec8c docs: add security reporting policy (#2561)
* docs(security): add private reporting policy

Document supported versions and route vulnerability reports through GitHub's private advisory workflow.

Add English and Chinese responsible-use notices to the READMEs.

Closes #2544

* ci: skip unrelated pull request builds

Use pull-request-aware path filtering for required Core, GUI, Mobile, and Test workflows so they still publish required check contexts without launching expensive jobs for documentation changes.

Limit the optional OHOS pull request workflow to relevant paths.
2026-09-10 12:30:49 +08:00
KKRainbow 993640b1dd fix(config): normalize [secure_mode] when loading TOML config files (#2562) 2026-09-10 08:56:57 +08:00