- rebuild: fold observations per fingerprint in a total order
(hcl DESC, node_id DESC, id DESC) so two nodes with identical logs
rebuild identical entries (was: arbitrary bare-column row, merge-order
dependent)
- threads: CreateThreadCluster uses INSERT OR IGNORE + existing-id
fallback — concurrent auto-threaders converge instead of hitting UNIQUE
- errors surfaced instead of swallowed: scanEntries returns rows.Err(),
Stats() fails fast on query errors, AutoLinkThreadObservations /
linkTemporalNeighbors / golden-thread linking propagate failures,
AddThreadNote + LinkObservationToThread write under one tx,
watch records session upsert failures
- gossip client: push checks HTTP status and reports errors (a broken
push direction no longer looks like a silent success); gossip diff
gets a 10s timeout so a dead peer cannot hang the CLI
- ingest: failed source sweeps (obsidian/browser/gitea) are logged,
and -d's help text now states its file-only scope
- watch --quiet: fatal errors go to stderr instead of io.Discard
- main: cobra SilenceErrors/SilenceUsage (errors print once, usage is
not dumped on runtime failures); knox mcp exits 0 on SIGINT/SIGTERM
- metrics: drop _total suffix from gauges (knox_observations,
knox_entries, knox_projects, knox_sessions, knox_peers, knox_threads);
_total stays on counters per Prometheus convention
- http: ReadHeaderTimeout + IdleTimeout on gossip, metrics, and web servers
tests: concurrent cluster-create idempotency, HCL-order rebuild fold
(both merge orders), push HTTP-error surfacing; full suite + -race pass,
gofmt clean
Implements the top findings from the codebase review, verified with tests and live CLI/MCP checks.
**Gossip integrity**
- Push validation: 4 MiB body cap, 1000-row batch cap; rows claiming the local node id (vector-poisoning), empty node ids, and negative HCLs rejected (internal/watch/gossip.go, internal/db/gossip.go)
- Reconcile-on-pull: Run returns the pulled count, syncGossip rebuilds derived state when > 0 — entry-count comparison could never fire, so synced observations never materialized into searchable entries
**Data-layer safety**
- HLC resumed from MAX(hcl) at Open (hlc.SeekTo): a restart with a regressed wall clock cannot reissue values the (node_id, hcl) locator and pull cursors depend on
- Writer serialization: _txlock=immediate DSN + SetMaxOpenConns(1) + per-KnoxDB mutex around RecordObservation's check-then-insert dedup (closes duplicate-row race)
**Watch daemon**
- Ticker guard flags now atomic.Bool (was a cross-goroutine data race)
- Trailing-edge per-path debounce (timer-based, pruned on fire/delete)
- Recursive watches (startup tree walk + watcher.Add on dir Create); Rename re-ingests, Remove cancels pending ingests
**MCP + CLI**
- Strict arg validation, no silent clamping: thread_id 0 errors instead of renaming thread #1; empty knox_thread_link {} errors instead of false success; thread existence checked before writes; golden-thread tool nil-safe
- --page 0 errors instead of panicking; query/recent pagination actually pages (page x limit)
**Tests** (new internal/hlc and internal/db packages): SeekTo monotonicity, concurrent dedup race, push validation, reopen HCL monotonicity, batch caps, self-spoof rejection, idempotency on observation counts.
Verified: go build, go vet, full suite with -race, live MCP stdio transcripts against a scratch DB.
Reviewed-on: #3
Co-authored-by: David Gwilliam <dhgwilliam@gmail.com>
Co-committed-by: David Gwilliam <dhgwilliam@gmail.com>
Refs #2
- /metrics served on a dedicated port (KNOX_METRICS_ADDR, default
localhost:8932) via prometheus/client_golang, with Go runtime +
process collectors
- DB-derived gauges refreshed per scrape: observations by source,
last-24h observations, entries, projects, sessions, pending
reflections, threads by status, peers, observations by origin node,
knowledge vector (max hcl per node)
- live gossip counters (pulls/pushes, observations pulled/pushed,
errors) incremented during the anti-entropy sweep; Run accepts an
optional metrics handle (nil for one-shot CLI)
- knox_node_info{node_id,name} for scrape identification
- internal/metrics package + db MetricsSnapshot; tests for snapshot,
scrape output, and counter increments
Refs #1
- /v1/diff endpoint returns each node's fingerprint set + tombstoned
auto-thread status
- knox gossip diff <peer-url>: shows peer-only/local-only fingerprints
(pull/keep preview) and tombstone divergence (resolved-on-peer vs
would-resurrect)
- knox gossip sync: one-shot anti-entropy sweep + reconcile
- gossip server starts before initial seed so peers can reach a booting
node; KNOX_PEER_ADDR alone now serves without KNOX_PEERS
- integration tests for diff + tombstone reporting
- e2e verified: two live daemons, diff previewed 1655 peer-only fps,
sync converged second node to 1655 observations