The readiness verdict was only reachable by attempting the migration, so the
UI's pre-flight panel could show every check green and then fail at Apply
with a 409. The user commits to the operation before learning it will be
refused.
GetMigrationSummary now runs the same read-only check and reports it:
data_ready_error carries the reason migration would be refused, and
data_ready_warnings carries the advisory ones. Nothing is enforced here, and
MigrateSpeaker still runs the check itself, so this cannot let a refused
migration through.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The GetExactDeviceInfo error was discarded and every failure reported as
"DeviceInfo.xml is not persisted", so a malformed file or an I/O error was
diagnosed as a missing sync and the user was told to run Data Sync, which is
the wrong remedy for either.
Include the cause, as the neighbouring branches already do.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
mapPresetsToFullResponse omits a preset whose source is absent from the
account's configured sources and cannot be synthesised. The readiness check
then found the speaker holding a slot /full does not, refused migration, and
attached its default action: "Run Data Sync for this device and retry
migration".
Syncing cannot add a missing music service source, so the user looped with no
override and no path forward.
compareMigrationPresets now reports whether the speaker's own view holds a
slot the rendered account lacks, which is the signature of that deliberate
omission, and that case gets an action naming the real remedy: re-link or
repopulate the source. Every other mismatch keeps the sync advice, which is
still right for a stale snapshot.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Clearing a slot through the Marge API leaves a zero-value entry in the list
(RemovePreset assigns models.ServicePreset{}), which savePresetsNoLock
persists as <preset id=""> with no filtering. Reading it back gives an empty
slot, while mapPresetsToFullResponse drops it from the rendered /full.
The comparison then refused migration with either "contains a preset without
a slot" or a count mismatch, for a datastore that was otherwise perfectly in
sync, and Data Sync could not fix it because nothing was actually wrong.
An empty slot carries no identity to compare, so skip it on both sides.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Migration was refused whenever the rendered account held more than one
device. One account holding every speaker in the household is the normal Bose
arrangement, so this blocked most setups, and there was no override: the
check runs unconditionally at the top of MigrateSpeaker.
It also misfired on genuinely single-speaker setups.
handleDiscoveredDeviceFallback writes a second device directory keyed by the
host address under the same account whenever /info momentarily fails, and its
cleanup only runs when d.SerialNo is set, which discovery never populates. A
stale entry left behind by a DHCP lease change then blocked migration
permanently.
The evidence does not support a hard block either. Issue #614 concluded the
shared-account preset wipe is empirical rather than a proven mechanism, with
the root cause still open, and the troubleshooting entry added there is
labelled a workaround. The guide's own remediation was unreachable in normal
use, since discovery re-adds the other devices.
The check now reports it as a warning, naming the device count, carried into
the migration log the UI already shows alongside its other "Warning:" lines.
The provable checks (persisted snapshot present and valid, presets equal
across snapshot, live /presets and rendered /full) still refuse.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The readiness check refused when the live /info carried no
margeAccountUUID, which is exactly the state a factory-reset speaker is in.
That made the documented onboarding impossible. The admin UI migrates first
and pairs afterwards (see "Pairing runs after the URL flip" in the setup
page), and MIGRATION-GUIDE.md step 4 tells the user to Generate an account ID
on a factory-reset device. Both now hit a 409 before pairing can run. The
suggested remedy could not help either: SyncDeviceData files an account-less
device under "default", which never matches an empty live account, so the
user had no way forward at all.
An unpaired speaker has no account data to preserve, so there is nothing for
this check to compare and nothing to lose. Skip it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
syncSources never reported how many sources it actually saved, so the
Admin UI's success message always said the meaningless "sources:
synced" regardless of outcome. syncSources now returns the count saved
(-1 if the fetch failed), threaded through SyncResult.SourcesCount.
Also replaces the single run-on results string (which visually mashed
presets/recents/sources together with no separator) with a real <ul>
list, one <li> per resource, matching the presets/recents diff lines.
Built via DOM APIs rather than innerHTML string concatenation, since
preset/recent names ultimately come from user-editable station names
on the speaker.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
SyncDeviceData's syncPresets/syncRecents unconditionally overwrote the
datastore with whatever the speaker's live :8090 API returned at that
instant, with no check against what's already stored. If the speaker's
own local cache was stale or incomplete at that moment (e.g. right
after a burst of preset writes, or shortly after a reboot before the
speaker resyncs with Marge), Sync would silently persist that bad
snapshot over good data. A reporter's fresh #614 repro showed the
account's /full response dropping from 6 to 5 presets right after a
Sync click, consistent with this mechanism.
SyncDeviceData now diffs a fresh live fetch against what's stored
before writing anything; if applying would shrink either list, it
returns the diff (via the new SyncResourceDiff/SyncResult types)
without writing unless the caller passes confirmed=true.
HandleInitialSync surfaces this as a 409 with the diff JSON; every call
(confirmed or not) re-fetches live from the speaker, so a confirmed
retry re-checks reality rather than replaying a stale snapshot. Sources
sync is left unconditional, as before -- lower risk in practice and
out of scope for this fix.
fetchLivePresets/fetchLiveRecents are extracted pure-fetch helpers;
syncPresets/syncRecents keep their unconditional-apply behavior (used
directly by existing tests) since the button-driven path now goes
through the diff/confirm guard instead.
Adds TestSyncDeviceData_DestructiveSyncRequiresConfirmation covering
both the refusal and the confirmed-retry path.
Frontend wiring (script.js's startSync + real per-resource result
rendering, replacing the current hardcoded "OK" text) is a follow-up
commit on this branch.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Speakers paired via non-AfterTouch tooling (e.g. the USB-stick SSH-enable
method) can report a margeAccountUUID that isn't Bose's own 7-digit
numeric format, such as "stick@local". Discovery persisted this value
unvalidated, and the datastore's identifier check rejected it outright,
so the device was silently never saved.
Widens datastore.IsSafeIdentifier to accept any identifier that's safe
as a path component, XML value, and telnet-command token (still
excluding whitespace, control characters, and HTML/XML/shell
metacharacters), and makes it the single account-ID validator,
replacing setup's separate, stricter 7-digit-only IsValidAccountID.
Also closes related gaps found while widening the validator:
- postSetMargeAccount now XML-escapes the account ID instead of raw
string interpolation.
- SaveAccountInfo/HandleMargeCreateAccount now validate the account ID
the same way SaveDeviceInfo already did.
- handlers_export.go URL-escapes account/device IDs before building
outbound diagnostic-fetch URLs.
- pkg/service/health gained the sanitizeLog helper every other package
already has, applied to log lines carrying speaker-reported values.
- The admin web UI (script.js) renders account/device IDs via DOM APIs
instead of innerHTML/inline event-handler string interpolation,
closing a stored-XSS path, and a duplicate escape helper was
consolidated into one.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A speaker can be reachable, named, and already account-paired yet still
report SOUNDTOUCH_NOT_CONFIGURED, leaving the "install the Bose app"
prompt on screen (reported for ST30 Series II/III in #615). Only a full
pass through the WebSocket setup state machine clears it, but running
that unconditionally risks re-running the bracket on speakers that
don't need or support it.
Add Manager.PreflightInitPlan: checks /supportedURLs for
/setMargeAccount, then requires /soundTouchConfigurationStatus to read
exactly SOUNDTOUCH_NOT_CONFIGURED before ExecuteInitPlan runs.
Already-configured devices are a no-op; an unsupported route or an
unrecognised status value aborts instead of guessing.
- setup: resync all four boseurls (not just marge/swUpdate) over telnet
after an SSH-XML migration. `envswitch boseurls set` persists whatever
is currently in the runtime layer, so leaving stats/bmx untouched froze
their stale pre-migration values into the persistence layer permanently
-- surviving reboot and previously requiring a factory reset to clear.
- admin-ui: Migrate tab's Target Domain edits now propagate into the four
service URL fields (tracked via a dataset.autofilled flag so real manual
edits still aren't clobbered), closing the gap where changing Target
Domain to a new value left the four fields pointed at a stale default.
- install.sh: prune stale binary backups before the download too, not
only after a successful install, so a backup left by a previously
aborted (out-of-space) run gets cleaned up instead of compounding.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Confirmed on real hardware (192.168.178.28): RevertMigration's full call
graph (revertXMLConfig/revertHosts/revertResolvConf/revertAftertouchHook/
removeRcLocalHooks/revertCACert) makes 17 separate client.Run() calls, and
pkg/ssh.Client.Run/UploadContent each dialed a brand-new SSH connection
per call with no reuse. Hitting a resource-constrained speaker with 17
rapid reconnects overwhelmed it -- confirmed via a follow-up plain SSH
command timing out at the TCP level, and the speaker going visibly
unresponsive.
Gives pkg/ssh.Client an opt-in persistent connection: Connect() dials
once and caches it, Close() releases it, and a shared dial() helper makes
Run/UploadContent reuse the cached connection when one's open, falling
back to today's per-call dial otherwise. RevertMigration now calls
Connect() once and defer Close(), collapsing 17 connections into 1. The
other ~21 m.NewSSH() call sites in pkg/service/setup never call Connect,
so their behavior is completely unchanged -- this only touches the one
function that was actually causing real-world problems.
SSHClient interface gained Connect()/Close(); both test mocks
(pkg/service/setup/setup_test.go, pkg/service/handlers/handlers_setup_test.go)
got no-op stubs. Added TestClose_NoOpWithoutConnect and
TestConnect_DialFailureLeavesConnNil in pkg/ssh/ssh_test.go -- these don't
prove connection reuse against a real server (Client.Run hardcodes :22,
no configurable port for a test listener), so that specific behavior is
verified by code review (a single `if c.conn != nil` branch) plus the
real-hardware confirmation above, not an automated integration test.
Also fixes the web UI's "Revert to Defaults" button, which calls the same
RevertMigration code path.
Addresses recommendations 4 and 7 from #515 comment 5231931569: a
green-looking getpdo readback only confirms the sys configuration
writes were accepted, not that they'll survive a reboot (that's what
the envswitch-persisted layer decides). Labels the getpdo line in both
migrateViaTelnet and runTelnetInjection's CLI/log output accordingly,
softens migrateViaTelnet's "succeeded" wording to "accepted", and adds
the same one-line caveat to TELNET-MIGRATION-METHOD.md #2.3 (previously
only in TELNET-COMMAND-REFERENCE.md).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Community hardware testing (bitranox, JRpersonal) on 2026-08-09 retracted
the earlier "inter-command delay is necessary" theory and established that
envswitch boseurls set commits the whole runtime layer (not just its two
arguments), has no read form, and doesn't ack with "OK". Corrects
TELNET-MIGRATION-METHOD.md and TELNET-COMMAND-REFERENCE.md accordingly,
retracts the stale "confirmed necessary" command-delay claim in
enable_ssh.go/cmd_setup.go, and lowers DefaultTelnetCommandDelay 5s -> 3s
as a smaller hedge now that the delay itself is known not to be the
mechanism.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Follow-up to the 3s default from earlier in #515: the reporter agreed
5s is a better trade-off (issue comment 5230881285) — more headroom
than the original guess, still comfortably under the ~7s gap their
manual A/B test used.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
#515 comment 5230833551: on a genuinely unpaired (factory-reset) device,
margeServerUrl is reportedly never polled at all, so the boseurls
SSH-enable injection has no read cycle to fire on regardless of any
command delay. enable-ssh now checks /info first and, if
margeAccountUUID is empty, pairs the device via the existing
PairAccount helper (HTTP /setMargeAccount, telnet fallback) before
running the injection.
Adds setup.Manager.EnsureMargeAccountPaired plus --no-auto-pair (skip
entirely) and --account (use a specific 7-digit ID instead of a
generated one, e.g. to match one already in the datastore) flags on
enable-ssh. Pairing failure is a warning, not fatal, since the claim
is unconfirmed on this specific hardware and existing working flows
must not regress.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Per #515 comment 5228449448: on a real Lifestyle console, the same 6
commands (5 sys configuration/envswitch + reboot) sent back-to-back left
sshd down after reboot, but succeeded sent one at a time with ~7s gaps —
same commands, same order, same device, minutes apart. Sending fast may
not let the device fully process one command before the next arrives.
Adds --command-delay (setup.DefaultTelnetCommandDelay, 3s), threaded
through EnableSSHViaTelnetFullConfig/runTelnetInjection (pause after each
of the 5 commands) and runEnableSSHInjection (one more pause before the
reboot). 0 restores the old back-to-back behavior. The reporter didn't
try to find the true minimum, just confirmed ~7s works and speculated
"a second or two may well be enough" — 3s is a middle ground, tunable via
the flag if a specific device needs more.
Also prints an approximate total for the injection phase up front (6
steps x delay, ~18s at the default) so the command doesn't look hung —
separate from the existing --wait message for sshd coming up after
reboot, which can take much longer.
Refs #515
checkCACertTrusted matched only the static "# AfterTouch" label in the
device's trust bundle. After the service CA was regenerated (e.g. a
recreated container with a fresh/empty data dir), the stale label was
still present, so the migration wrongly reported the speaker as already
trusting the new CA and skipped re-installing it, leaving the speaker
unable to validate TLS to the service.
When the service CA is available, compare the actual cert payload and
re-install on mismatch; fall back to the label only when the CA can't be
read (CLI callers without Crypto). Adds regression tests for the
stale-label and no-Crypto cases.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The default `setup enable-ssh` injects the remote_services/sshd payload
only via `envswitch boseurls set` and relies on the speaker re-reading its
boseurls (~60s) without a reboot. On the SoundTouch Portable (Series I,
FW 27.0.6.46330.5043500) and some CineMate 520 units the device accepts and
persists that injection (getpdo confirms) but sshd never comes up, so :22
stays "Connection refused".
@Henri-be got root on the ST Portable by typing a different sequence by hand
over telnet :17000: the injection rides `sys configuration margeServerUrl`
(the runtime layer) as well as `envswitch`, all four URL keys are written,
and the device is rebooted so it re-parses the config at boot.
Add an opt-in `--full-config` flag that replicates that exact sequence
(EnableSSHViaTelnetFullConfig + telnet reboot via the existing
RebootMethodTelnet). The default single-envswitch path is unchanged, so the
field-confirmed flow on the Wireless Link Adapter and CineMate 520 `lisa`
variant does not regress. Docs (TELNET-COMMAND-REFERENCE, DEVICE-LOGGING)
document both paths and which device models/firmware need `--full-config`.
The flag automation is candidate behaviour awaiting reporter confirmation:
the manual sequence is confirmed on the ST Portable, the flag is not yet.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two follow-ups from the #471 field reports on the BETA `setup enable-ssh`:
1. enable-ssh: when sshd (:22) does not come up within the wait window, this is
no longer treated as a hard error. On some devices (e.g. the Wireless Link
Adapter) the envswitch injection is accepted but sshd only starts after the
speaker restarts. The command now prints a warning with power-cycle + retry
guidance (and the exact ssh command), deliberately leaves the injected
boseurls in place so a restart re-triggers the unlock, and exits cleanly
instead of failing.
2. XML migration: re-apply the boseurls over telnet at the end of migrateViaXML
so the runtime layer reported by `getpdo CurrentSystemConfiguration` matches
the persisted SoundTouchSdkPrivateCfg.xml. After enable-ssh bootstraps SSH,
that runtime layer still points at the placeholder (https://aftertouch.invalid),
so the preflight cross-check keeps warning that margeServerUrl/swUpdateUrl
differ between transports until a reboot. The re-apply reconciles it now.
Best-effort: if telnet is unavailable (e.g. port 17000 was closed via
--close-17000), a reboot still reconciles the layers, so it only logs a note
and never fails the migration.
Tests cover the re-apply command and its best-effort (non-fatal) behavior.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds the #471 "secure" steps as opt-in flags on `setup enable-ssh`, off by
default (per the decision that closing 17000 must be opt-in):
- --close-17000: blocks port 17000 from the LAN. Manager.Close17000 remounts /
read-write, persists an idempotent iptables rule in
/etc/init.d/Firewalls/update_iptables (keyed on a marker), and applies it
immediately; loopback access is kept.
- --authorized-key <pubkey>: Manager.InstallAuthorizedKey writes the key to
/home/root/.ssh/authorized_keys so root SSH no longer relies on the
empty-password login.
Both run over the SSH the enable step just opened. Default output reminds the
user that 17000 is left open and how to close it. Unit tests cover the
firewall command sequence and the key upload path.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds `soundtouch-cli setup enable-ssh`, the first iteration of foob61451's #471:
turn on SSH on a speaker that has no prior SSH access and without a USB
recovery stick, then fall into the migration / CA-install flow we already have.
Mechanism (new Manager methods, reusing the existing telnet :17000 client):
- EnableSSHViaTelnet sends `envswitch boseurls set "<url>;touch
/tmp/remote_services;/etc/init.d/sshd start" "<url>/update"`. The injected
shell commands run when the speaker next parses its boseurls (~60s), starting
sshd. The URL is only the vehicle for the injection — it does NOT need a live
server, so this works before any AfterTouch service exists.
- WaitForSSHPort polls :22 until sshd is up.
- ResetBoseURLs restores a clean marge URL afterwards.
- Persistence reuses the existing EnsureRemoteServices (writes the marker over
the now-open SSH so it survives reboot).
CLI flow: inject → wait for :22 → reset clean URLs → persist. `--service-url`
is optional (placeholder used otherwise; set real URLs later via migration).
Securing/closing port 17000 is deliberately OPT-IN and not done here. Unit
tests pin the exact injected/reset command strings and the double-quote guard.
This lands in -cli first (cheapest to iterate); the future soundtouch-app can
reuse the same Manager methods.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Manager.HTTPGet defaulted to http.Get, which uses http.DefaultClient with
no timeout. An offline speaker therefore hung the caller for the OS-level
TCP timeout (~30 s). The admin device list refreshes every device's live
/info on each load (updateDeviceInfo per row), so a handful of offline
speakers each held a request for 30 s. Server-side those run concurrently
and never blocked other routes, but the browser's ~6-connections-per-origin
limit got saturated by the long-held /info requests, which made the whole
admin page (and navigating away from it) feel stuck.
Give HTTPGet a 5 s timeout (liveDeviceHTTPTimeout): ample for a healthy
speaker on the LAN, quick to fail a dead one. Applies to the /info,
/presets, /recents, /sources, inspect, and peer-probe GETs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Per the repo's no-real-data rule (CLAUDE.md), scrub committed files only (the
gitignored _/ local captures are left as-is):
- Real Bose-OUI device ID 08DF1F0BA325 -> placeholder AABBCCDDEE0A across 4 docs
and 8 Go test files (consistent 1:1 rename; affected packages tested green).
- Personal/topology LAN IPs -> RFC-5737: the lab runbook's AP subnet
192.168.10.x -> 198.51.100.x (192.0.2.x is already used contrastively there)
and 192.168.100.1 -> 203.0.113.1; illustrative example IPs in
ANONYMIZATION-SUMMARY / spotify-overview / TROUBLESHOOTING -> 192.0.2.x.
- Kept factual RFC-1918 range citations (10.0.0.0/8 trusted-proxy example,
192.168.0.0/16 "all private subnets") since they name the ranges themselves.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
soundcork#104 confirms speakers validate the /speaker audio-notification
app_key against audionotification.api.bosecm.com (100 calls/day on real
Bose). Our /v1/auth shim accepts it, but a host-seeded migration only
worked if the speaker resolved that host to us. DNS interception already
covers it (bosecm.com substring), but the /etc/hosts migration domain
list did not — so the speaker method would fail on hosts-based setups.
Seed both audionotification.api.bosecm.com and the dev variant
(audionotificationdev.api.bosecm.com; firmware may use either) into the
migration /etc/hosts lists, and update the mock fixtures/docs accordingly.
/v1/auth is path-based, so it already answers regardless of which host the
speaker thinks it is calling.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The diagnostic export captured the symptom of #345 (a TuneIn select
escaping to the dead Bose Apigee gateway → BMX_HTTP_ERROR 4501 →
INVALID_SOURCE) but none of the data that decides where a speaker sends
its marge/BMX/streaming traffic, so we couldn't tell whether the request
was ever redirected to AfterTouch.
Collect that per speaker:
- New collectSpeakerRedirectConfig prefers the on-device
SoundTouchSdkPrivateCfg.xml over SSH (archives raw + parses
marge/stats/swUpdate/bmxRegistry URLs), and falls back to
`getpdo CurrentSystemConfiguration` over telnet when SSH is
unavailable — the same channel the telnet migration uses. Parsed URLs
and provenance land in diagnostic.json as redirect_config: source
(ssh|telnet|none), ssh_reachable, and inferred_migration_method
(telnet when only telnet answered, since xml/hosts/resolv all need SSH).
- Pull redirection-relevant files over SSH: /etc/hosts(.original),
/etc/resolv.conf, the resolv-method hook, /mnt/nv/remote_services, and
the pre-migration .original backups (CA bundle and the URL config).
- Dump the speaker firewall (iptables-save; ip6tables-save is empty on
FW 27.0.6 but harmless) to catch self-inflicted DROP rules (cf. #354).
Export ParseGetpdoConfig from pkg/service/setup and add a test pinning
the field-name contract the export depends on.
Diagnostic-collection only; does not change migration or playback.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Root cause of #334's INVALID_SOURCE: a speaker reports device-local slots
(STORED_MUSIC_MEDIA_RENDERER, UPNP) in /sources; AfterTouch imports them
verbatim and re-serves them in /full. PrepareConfiguredSource fills
sourceproviderid only for types in constants.StaticProviders, so these go
out with an empty <sourceproviderid> — a required protobuf field — and the
speaker rejects them as INVALID_SOURCE, which then re-syncs back into the
datastore.
Fix, keyed on the principle (no hardcoded denylist in production):
- HasResolvableProviderID(s): true if the source already carries a provider
id, or its source-key type resolves via StaticProviders.
- Serve-side guard in getAccountSources: drop any source whose resolved
sourceproviderid is still empty (generalises the existing AUX/#195 skip).
Heals already-polluted datastores on the next /full, no resync needed.
- Import-side filter in syncConfiguredSources (marge) and both branches of
syncSources (setup): drop unresolvable sources before persisting, stopping
future pollution and the re-import loop.
Tests: reproduction converted to regression test
(TestI334FullOmitsSourcesWithoutProviderID) seeded from a sanitised real
#334 /sources capture; explicit servable/non-servable tables in
TestHasResolvableProviderID. Two pre-existing fixtures that relied on
sources with no provider id were given valid ones.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two per-device checks run against each speaker's CA bundle via a
single SSH probe round-trip:
(1) Every PEM block from ca-bundle.crt.original (the factory backup
written by TrustCACertFromBytes on first CA injection) must be
present in the live ca-bundle.crt. A missing block means the
original trust store was truncated, which would break external
HTTPS (Spotify, Amazon, firmware updates).
(2) The AfterTouch CA sentinel (# AfterTouch) must be present in
the live bundle. Without it the speaker rejects AfterTouch's
TLS cert and migration is effectively inactive.
Both findings carry a QuickFix:
- FixIDRestoreAndInjectCA: cp .original → live bundle over SSH,
then TrustCACert to re-inject the AfterTouch CA.
- FixIDInjectCACert: TrustCACert only (original certs intact).
Graceful degradation:
- SSH unavailable → SeverityInfo, no fix offered.
- .original absent (device never had install-ca run) → SeverityWarning,
suggest install-ca; check (2) still runs.
Infrastructure changes:
- ssh_probe.go: add ca-bundle.crt.original to probeFilePaths (free
in the existing single-round-trip batch).
- setup.go: export ProbeCABundles and RestoreCABundleFromOriginal so
the handlers package can use them without exposing speakerProbe.
- Fix executors live in handlers (need setup.Manager) per the
established boundary used by completeSpeakerPairingFix.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- stockholm/static.go: wrap deferred root.Close() in func(){}() to
silence errcheck; change 'rel = rel + ...' to 'rel += ...' (gocritic).
- Remove sanitizeErr from four logutil files where no call site exists
(cmd/soundtouch-cli, cmd/websocket-demo, pkg/discovery, pkg/service/setup).
The log-injection fixes in those packages used sanitizeLog on string
arguments rather than sanitizeErr on error values.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- isXMLMigrated and isResolvConfMigrated now guard against empty hostname
(Go's strings.Contains(s, "") is always true, causing any speaker to
appear migrated when --service-url has a malformed single-slash scheme)
- renderPlanSteps message no longer claims "and paired" when --include-pair=false
- validateServiceURL rejects malformed service URLs early with a hint
(e.g. "did you mean https://soundtouch.fritz.box?")
- Generated plan-step commands move --host before the subcommand name
(urfave/cli/v2 requires global flags before the first subcommand token)
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Mirrors the .md/.txt sweep across all tracked _test.go, testdata XML,
and .http integration files. Test files are self-contained (producer
+ assertion in the same file), so the matched-pair swap stays green
under `go test ./...`.
Mapping applied:
192.168.178.[0-9]+ → 192.0.2.[same]
192.168.1.[0-9]+ → 192.0.2.[same]
Sound Machinechen → Living Room SoundTouch
A Sound Machine → Kitchen SoundTouch
A81B6A536A98 + case/separator variants → AABBCCDDEEFF (etc.)
A81B6A849D99 → AABBCCDDEE01
A81B6A849D88 → AABBCCDDEE03
A81B6A536A09 → AABBCCDDEE04
884AEAEEBD27 → AABBCCDDEE02
3230304 → 1000001
9569497 → 1000002
Two semantic fixes alongside the bulk swap:
- pkg/service/zeroconf/zeroconf_test.go: the "private 192" and
"strips query" cases pin acceptance of RFC-1918 192.168/16. They
must use a real 192.168 value; doc-range IPs would (correctly) be
rejected by validateZcBaseURL. Switched to 192.168.10.10 — generic
enough not to match any home LAN default, real enough for the
validator. Added a comment explaining why this single test still
carries a 192.168 literal.
- pkg/service/setup/setup_test.go: TestTestDNSRedirection mocks the
device's `od -An -tu1` byte output, which is space-separated
octets ("192 168 1 100"). My sed only matched the dot-separated
form, so the mock was returning the old IP while the test
assertions had moved to the doc range. Updated to " 192 0 2 100".
go build ./... clean. go test ./... clean (only TestDocsConsistency
remains failing, which is a pre-existing/untracked-file issue).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The ST10's /presets response after a factory reset emits self-closing
<preset/> entries with no ContentItem child. cmd/soundtouch-cli's
getPresets() handled the missing ContentItem in GetDisplayName() but
then dereferenced preset.ContentItem.Source on the next line, panicking
with "invalid memory address or nil pointer dereference" the moment the
loop reached the first empty entry.
A second placeholder shape was observed on healthy devices that were
never reset: <preset id="0"><ContentItem source="INVALID_SOURCE"
isPresetable="true"/></preset>. ContentItem is non-nil here, so the
previous "ContentItem != nil" guard at other call sites still let
these placeholders through into listings and into the AfterTouch
datastore.
Fix shape:
pkg/models/presets.go - extend Preset.IsEmpty() to recognise both
shapes (ContentItem == nil, OR Source == "" / "INVALID_SOURCE").
HasPresets, GetEmptyPresetSlots and GetUsedPresetSlots become honest
about which slots actually carry playable content.
cmd/soundtouch-cli/cmd_info.go (the crash site) - filter the slice
via IsEmpty before the print loop, and switch the still-printed
fields to the existing nil-safe Get* helpers.
pkg/service/setup/setup.go - upgrade syncPresets's "ContentItem ==
nil" continue-guard to IsEmpty so Shape B placeholders don't get
persisted in the AfterTouch datastore and then surface as junk
rows in the admin web UI.
cmd/soundtouch-cli/cmd_events.go, cmd/websocket-demo/main.go - same
nil-guard upgrade. These already nil-checked so were crash-safe;
the change is for consistency and to stop printing
"Preset 0: (INVALID_SOURCE)" demo lines.
examples/preset-management/main.go - had the same latent crash as
cmd_info.go; same fix shape.
Regression tests in pkg/models/presets_test.go cover both shapes using
the exact XML observed in the wild: the reporter's three <preset/>
placeholders plus the three INVALID_SOURCE entries from a live device.
The reporter XML test walks every preset through the same accessor
path the CLI used and asserts no panic.
The soundtouch-web Go code does not deref preset.ContentItem.X
anywhere - presets flow through as JSON - so no separate crash trap
exists there. The web frontend will pick up the cleaner data once
syncPresets stops persisting placeholders.
Closes#308
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two doc-only additions to TestValidateRealSpeakerBundle's header
comment:
- Cross-model note: ST10 and ST20 ship the byte-identical CA
bundle on firmware 27.0.6.46330.5043500 (md5
2d150987b312e4280fc576b508e62b43, 165 certs, ~251 KB).
Verified against firmware/_backup_ST10/_/etc/pki/tls/certs/
ca-bundle.crt 2026-05-16. The existing
testdata/ca_bundle_st20_pristine.crt fixture therefore stands
in for both models on that firmware build, so any expired-root
hypothesis evaluated against it covers both.
- Curl reproducer: three one-liners that point curl at the fixture
and probe the actual TuneIn stream chain a SoundTouch speaker
would walk (using K-LOVE / s33828 as the canonical example —
matches the case from #292). Control with the system trust
store shown alongside. Both bundles handle the chain (Amazon
Root CA 1 + DigiCert Global Root, valid through 2026+) so the
expired-root hypothesis is ruled out for firmware 27 — recorded
in the comment so future-me / reviewers can replay the same
probe without re-deriving it from chat context.
No code change; test still passes.
Related to https://github.com/gesellix/Bose-SoundTouch/issues/292.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The migration-summary preflight always emitted a "resolved from service,
not from device" ❌ row whenever the target was a hostname — even when
SSH was available and could have answered authoritatively. Two
problems compounded: the summary builder passed `nil` for the SSH
client (skipping the device-side ping), and resolveIP's service-side
fallback returned a bare fmt.Errorf the caller couldn't distinguish
from a real failure.
Changes:
- ErrResolvedFromServiceOnly sentinel; service-side fallback wraps
it with fmt.Errorf("%w: ...") so callers can errors.Is()-check.
Apply-path callers that pass a real SSH client keep getting the
same error shape they always did.
- populatePlannedNetworkConfig now takes an SSHClient. GetMigrationSummary
opens one when probe.SSHOK is true and passes it through, so the
summary's resolve call uses the same device-side authority the
apply paths use. Skipping the dial when SSH is known dead keeps
a stale handshake-timeout from burning the preflight budget.
- MigrationSummary gains ResolveIPSource ("device" / "service") and
ResolveIPDurationMS so we can observe the SSH-ping cost in the
wild. The historical comment claimed 2-5 s on firmware-27 devices —
we now have data instead of a guess.
- CLI renderer prints the new source + timing line, and only renders
the ❌ ResolveIPError row for hard failures (both SSH ping AND
service DNS failed).
- Two regression tests cover the sentinel-tagging contract and the
device-success-returns-nil-error path.
Related to https://github.com/gesellix/Bose-SoundTouch/issues/282.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Pins the ordering invariant fixed in the preceding commit. Builds a
fake-speaker scenario where:
- SSH is unavailable (every SSH-driven axis stays false)
- telnet getpdo reports the AfterTouch hostname
Pre-fix, checkIsMigratedFromProbe ran before the telnet channel was
drained, so summary.TelnetVerifiedConfig was empty when
isTelnetMigrated read it — the telnet axis came back false and
summary.IsMigrated followed. The CLI's `setup verify` exited
non-zero, the web UI rendered "Not Migrated". Reproduced by
foob61451 on #293.
The test asserts:
- summary.TelnetVerifiedConfig is populated (sanity guard — the
downstream assertions are meaningless if the probe didn't run)
- summary.TelnetMigrated == true
- summary.IsMigrated == true
Verified locally: the test PASSES with the ordering fix applied and
FAILS without it. Failure messages name PR #294 by number so a
future regression points at the same code path.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The previous 10s→30s timeout bump didn't help — the first POST to
/addWirelessProfile on the speaker's AP-mode endpoint frequently
hangs until the deadline elapses, then a second POST a few seconds
later succeeds immediately. Empirically the workaround was "just
run wifi-push twice"; this commit folds that into the function.
PushWiFiCredentials now:
- caps each attempt at 12 s (well above the sub-second healthy
response time) so a stuck first attempt doesn't burn the whole
budget
- waits 2 s between attempts so the speaker's setup endpoint can
finish whatever the first POST kicked off
- falls through cleanly if the first attempt succeeds (the second
never fires)
- returns the second attempt's error if both fail, with context
cancellation surfaced explicitly
Total budget is well under the CLI's 30 s --request-timeout, so
the flag still acts as a hard ceiling.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The Stockholm app (stockholm/setup/js/workflow_add_devices.js:23,77)
and Zimbo88's OpenCloudTouch USB-less script
(https://github.com/scheilch/opencloudtouch/discussions/201) both send
<boseServer>, <updateServer>, and <accountEmail> alongside the
<accountId>/<userAuthToken> pair. AfterTouch's setMargeAccount
historically sent only the latter two.
Adds:
- MargePairingExtras struct on SessionConfig, opt-in via
BoseServer (UpdateServer + AccountEmail default-derived when
empty).
- DefaultMargeAuthToken constant ("Bearer AfterTouch") and
DefaultMargePairingEmail constant ("local@aftertouch.invalid",
RFC 2606 reserved .invalid TLD).
- buildPairDeviceWithAccountXML helper extracted so tests can
pin both the minimal-payload and extended-payload shapes
without driving a full WebSocket session.
- --token flag on `soundtouch-cli setup pair` so we can override
the placeholder for token-shape experiments.
- runPairBare threads --service-url through to PairingExtras so
`--mode=bare --service-url=...` ships the extended payload too;
runPairFull already used it via applyInitPlanDefaults.
The speaker accepts any non-empty Bearer string (verified during
#195 investigation: "Bearer AfterTouch" passes and the speaker
re-derives its post-pair state from the marge endpoints regardless
of token content). The Stockholm-app payload shape is purely
documentation alignment; it did NOT fix the post-pair AUX/preset
breakage that turned out to be the cloud /full source list (see the
preceding marge commit). Keeping the wiring so the switches are
ready when we want to experiment further.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The speaker confirms AddWirelessProfile then tears down its AP within
~30 s. The default 10 s --request-timeout races that ACK whenever the
speaker is busy reconciling state — and a hard-coded 10 s on the
internal http.Client capped the user-passed timeout silently, so a
longer --request-timeout had no effect.
The CLI default is now 30 s and the inner http.Client lets the
context govern alone.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes the AfterTouch-side half of issue #234. After a factory reset
the speaker's /sources only lists the always-on local entries (AUX,
BLUETOOTH, AIRPLAY, NOTIFICATION, QPLAY, plus a SpotifyConnectUserName
placeholder); TUNEIN, LOCAL_INTERNET_RADIO, DEEZER, and linked
Spotify accounts are absent until the device receives the
<sourcesUpdated/> notification the reporter ran by hand. SyncDeviceData
now POSTs that notification as the final step, so users get the
visible-source-list recovery for free when they click Data Sync.
The other half — re-creating Marge.xml so playback resumes — is
already handled by the wizard's pair-account flow: it detects an
empty <margeAccountUUID/> in /info and prompts the user to pick a
known account or generate a new one. The wizard's pairing UI is
deliberately user-driven (the user picks the ID); the notification
nudge is purely automatic because there's no choice to make.
Implementation routes through the existing client surface rather
than reinventing it. setup.notifySpeakerSourcesUpdated delegates to
pkg/client.Client.NotifySourcesUpdated — the same path
handlers_mgmt.go already uses after music-service account changes
(handlers_mgmt.go:304, :637). The wire shape lives in one place
(pkg/models.NewSourcesUpdatedNotification). Fire-and-forget: a
notification failure logs but doesn't fail the sync.
Adjacent UX changes:
- docs/guides/TROUBLESHOOTING.md: new section "Presets flash then
revert to 'Select a preset' after a factory reset". Names the
symptom, the Marge.xml + reduced-/sources cause, and walks the
user through re-opening the Migration tab + Data Sync.
- pkg/service/handlers/web/js/script.js: devices list now renders
a "⚠ Not paired — re-pair" badge in the account-ID column for
speakers whose live /info reports an empty margeAccountUUID.
Clicking it opens the Migration tab pre-filled with that device,
surfacing the wizard's existing "Not paired (factory-reset or
never paired)" flow without making users discover it cold.
- pkg/service/testing/fakespeaker/testdata/info.xml: demo speaker
now reports margeAccountUUID=1234567 instead of the misleading
0000000 (which AfterTouch happens to accept as syntactically
valid but is not a documented sentinel anywhere — the convention
is empty for factory-reset, a real 7-digit number otherwise,
matching pkg/client/testdata/info_response_st{10,20}.xml).
Screenshots regenerated accordingly.
Test scaffolding:
- fakespeaker grows a POST /notification recorder that captures
body + Content-Type; tests assert on s.Notifications().
- TestIssue234_FactoryResetSpeakerSyncsReducedSources now drives
SyncDeviceData end-to-end (exercises the wiring) and asserts
the notification fires with the right deviceID and shape.
- TestFakeSpeakerNotificationRecorder pins the recorder contract
and the POST-only method gate.
Refs #234.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Hardens TrustCACertFromBytes against the failure mode behind issue
#262 (corrupted /etc/pki/tls/certs/ca-bundle.crt on a SoundTouch 20)
and against silent transport-time corruption of our own writes.
Three-part change.
1. Atomic write path. The previous flow piped bytes straight into the
live bundle via `cat > <path>`; a dropped SSH session or partial
write left the device with a half-written trust store and no way
to roll back. The new path:
- uploads to <bundlePath>.aftertouch.tmp (sibling on the same
filesystem, same rw remount),
- reads the tmp back over SSH,
- validates the readback at the PEM-frame layer + the AfterTouch
sentinel bracketing,
- atomically `mv`s the tmp into place,
- on any verification failure: `rm -f` the tmp; the live bundle
is never touched, so there is no rollback semantics to reason
about.
The .original backup written on first install stays as
defense-in-depth (manual recovery for corruption from outside this
code path), but it is no longer the primary safety net.
2. New validators in pkg/service/setup/ca_validation.go.
- validateCABundleBytes: BEGIN/END marker counts match, every
decoded block is a CERTIFICATE with a non-empty body, decoded
block count equals BEGIN-marker count (catches a block with
unparseable base64 body), trailing non-PEM/non-comment content
rejected.
- validateAfterTouchLabelBracketing: CALabel appears exactly
twice and brackets exactly one CERTIFICATE block.
- stripAfterTouchEntries: collapses any number of stale
AfterTouch entries from the existing bundle. Older releases
reported to have appended without stripping, so long-lived
devices can carry several copies; we strip them all and log
the cleanup count rather than failing validation. Unpaired
sentinels (truncated prior install) surface as a structured
anomaly the caller logs and warns about.
The validators stay at the PEM-frame layer on purpose — an
earlier iteration called x509.ParseCertificate per block and
rejected the real ST20 bundle on block 29 (Go 1.23+ disallows
negative serial numbers, but Mozilla CCADB still ships ancient
CA roots that have them). Shipping that version would have made
every legitimate speaker install fail. The corruption mode #262
surfaces at the PEM-framing layer; x509-level checks aren't what
we needed.
3. testdata/ca_bundle_st20_pristine.crt is the pristine
/etc/pki/tls/certs/ca-bundle.crt captured off a real SoundTouch 20
(firmware 27.0.6.46330.5043500, snapshot 2022-08-04). Mozilla
CCADB public dataset, 165 certs, ~251 KB. TestValidateRealSpeakerBundle
locks in the cert count and asserts the strip pass is a no-op
against a bundle that has never been touched by AfterTouch.
Test infrastructure. mockSSH (both the setup-package and the
handlers-package copies) now mirrors UploadContent into a private
map so a subsequent `cat <path>` on the same path returns what was
written there. Lets the tmp-readback step in TrustCACertFromBytes
work against tests that only scripted the live-bundle path, without
per-test wiring. Two new behavioural tests in setup_test.go:
TestTrustCACert_StripsMultipleStaleEntriesSilently (pins the
multi-entry cleanup contract) and
TestTrustCACert_PostUploadVerificationFailureCleansUpTmp (pins the
rollback-free recovery: live bundle untouched, tmp removed).
Refs #262.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two-part iteration. First, the fakespeaker grows a `/now_playing`
route with a default STANDBY fixture — issue #235 is the first one in
this series that needs to override /now_playing, and adding the route
on its own would be infrastructure noise; bundled here it has an
immediate consumer.
The regression test then locks in the device-side signal at the heart
of #235: when a SoundTouch is targeted by Spotify Connect (Spotify
app sends audio to the speaker), the speaker's /now_playing reports
- source = SPOTIFY
- sourceAccount = SpotifyConnectUserName (the marker)
- ContentItem.location = /playback/container/<base64 spotify:...>
— a perfectly resolvable URI
- **ContentItem.isPresetable = false**
The contradiction (resolvable location + isPresetable=false) is the
reason the CLI's storeCurrentPreset at
cmd/soundtouch-cli/cmd_preset.go:41 refuses to act and emits "current
content cannot be preset" — exactly the reporter's symptom.
The test base64-decodes the location to surface the contradiction
explicitly: it should yield a `spotify:` URI. When AfterTouch grows a
fallback path (CLI --force, or service-side resolution to the
device's own Spotify integration via the SoundTouch Spotify source
provider), the assertion here stays sound — it tests what the device
emits, not what the CLI decides — but a sibling test should assert
the new fallback path produces a successful preset.
Fixture pattern matches the rest of the issue series:
testdata/issue235/ next to the test, fakespeaker driven via
FixtureOverrides, doc-comment naming what would have to change for
the assertion to flip.
Refs #235.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Wires the device-side state the reporter described in
https://github.com/gesellix/Bose-SoundTouch/issues/234 into the
fakespeaker via FixtureOverrides, and exercises GetLiveDeviceInfo +
syncSources against it.
The factory-reset state has two observable signals:
- `/info` returns an empty `<margeAccountUUID/>` because Marge.xml
is missing from the persistence partition. AfterTouch's
"is the device paired?" check at setup.go:632 keys on AccountID,
so this is the canonical "needs re-pairing" signal.
- `/sources` lists only AUX, BLUETOOTH, AIRPLAY, the
SpotifyConnectUserName placeholder, NOTIFICATION, and QPLAY —
TUNEIN, LOCAL_INTERNET_RADIO, and any post-pairing Spotify
accounts are gone until the speaker is nudged with a
`<sourcesUpdated/>` notification or re-pairs.
Today AfterTouch has no auto-recovery for either signal — it just
passes the state through. The test locks in that contract by
asserting:
- GetLiveDeviceInfo reports an empty MargeAccountUUID,
- persisted Sources.xml contains AUX/BLUETOOTH/AIRPLAY sourceKeys,
- persisted Sources.xml does NOT contain TUNEIN/LOCAL_INTERNET_RADIO.
When auto-recovery lands (e.g. an automatic POST of the
sourcesUpdated notification during sync, or marge-side source
replenishment), the absence assertions will flip — at which point
update them to assert the survivors are *present*, and adjust the
doc-comment so the contract stays in sync with the code.
Pattern mirrors pkg/service/setup/issue218_regression_test.go: a
testdata fixture next to the test, fakespeaker driven via
Config.FixtureOverrides, doc-comment naming what would have to
change for the assertion to flip.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>