Separate the update decisions
Follow five separate questions: who authorized the update, which devices may install it, whether the package is intact and current, how activation survives interruption, and what recovery does when a check fails.
The plates connect those decisions in one running example. You do not need to know a particular update standard in advance; each technical term is introduced where it changes the device's decision.
A water station needs a security fix—what could go wrong?
A water-level station sits far from the maintenance team. Its vendor has fixed a vulnerability, so the station needs new firmware. But the update arrives over a network the operator does not fully control, and power may fail halfway through. An attacker could also replay an older, correctly signed release that still contains the flaw.
The asset is not just the firmware file. It is the station's ability to report trustworthy readings and the signing authority that decides which software may run. The attacker may control the delivery path or possess an old, valid package, but we are not assuming they can break the device's protected trust anchor. That boundary matters: no update protocol can repair a trust anchor that has already been fully compromised.
The device therefore has to answer several separate questions before replacing its working software: who authorized this package, is it meant for this exact device, is it intact, is it allowed under rollback policy, and can the device recover if installation is interrupted? A valid signature answers only part of that list. [1][2][3]
The download server delivered it. Does that make it authorized?
A server can deliver bytes without having the authority to approve them. A mirror, cellular gateway, or compromised update server might send a package the vendor never intended this model to install. Encryption of the network connection can protect a conversation in transit, but it does not by itself establish that the firmware publisher approved the image.
Instead, the device needs a local trust path to an authorized signer or update authority. A signed manifest can state which image is being offered, which device class may accept it, and what policy fields apply. The device verifies that signature using trust material provisioned or updated through a protected process; it should not accept a new signing key merely because the download server names it. [1]
This is the distinction between delivery and authorization. The network helps move a candidate package. The device's own policy decides whether that candidate has an approved source and scope. If the verification key or policy can be replaced through the same untrusted path, the check collapses into trusting the attacker who supplied the bytes. [1][2]
A genuine package can still be for the wrong device
Suppose the vendor really did sign a firmware image—but for a different board revision. The signature can be perfectly valid and the image can still be unsafe for this station. Hardware revisions may differ in memory layout, peripherals, boot configuration, or security capabilities; installing the wrong build could leave the device unusable or weaken its protections.
That is why update authorization needs a target check, not just a publisher check. A manifest can identify the vendor, product class, component, hardware compatibility, and dependencies. The device compares those fields with its own protected identity and current configuration. A broad family name may not be precise enough if two models have different hardware.
Compatibility is not a cosmetic label. It connects the signed release to the specific installation decision being made. Product teams must define which identifiers are trustworthy, how revisions are represented, and what to do when the device cannot reliably establish its own target identity. A check is only as good as the identity and policy behind it. [2]
A valid signature can still carry an old vulnerability
Now imagine the station already runs the repaired release. An attacker sends an older package that was genuinely signed before the vulnerability was discovered. The signature still verifies: it proves that an authorized key signed those bytes, not that the release remains acceptable today.
The device needs an anti-replay or rollback policy in addition to signature verification. One design uses a monotonically increasing sequence number in the signed manifest and remembers the newest accepted value in protected storage. If the station has accepted sequence 112, a package with sequence 108 is rejected—even if its signature is authentic and its firmware version looks familiar. [2]
Do not confuse that sequence number with the human-readable firmware version. RFC 9124 Section 4.3.1 explicitly permits a newer manifest sequence number to authorize a return to an older firmware version. The rule is not simply ‘larger version string wins’; it is ‘accept only releases authorized by the device's current update policy.’ [2]
What exactly should the device check before installation?
A sound decision is a series of checks, each answering a different question. Verify the manifest's signer and authority; confirm the target and dependencies; check the image digest against the signed manifest; apply sequence and rollback policy; and confirm the device has enough resources and a supported installation path. A single green signature indicator cannot stand in for all of these decisions. [1][2]
The digest links the manifest to the exact image bytes. If one byte changes after the publisher calculated the digest, the recomputed value should differ. The signature protects the manifest's claims against unauthorized alteration, while the digest lets the device detect whether the fetched image matches those claims. Neither check tells the device whether the release is compatible or operationally safe—that comes from target policy and testing. [2]
The checks also need to apply to the bytes that will actually be installed. If the device verifies one copy and later reads a different, mutable copy, an attacker may exploit the gap between checking and use. Robust designs bind verification to the staged image and protect that image and its metadata until activation. [1][2]
Why not overwrite the only working image directly?
If the station erases its only firmware image and then loses power, it may have nothing left to boot. The update could be authentic and correctly targeted, yet the device would still become unavailable because installation did not finish. Remote locations make this failure especially costly: a technician may need to travel simply to restore service. [1][3]
A safer pattern is to keep the current working image intact while writing the candidate somewhere else. The device stages the new image, verifies it, and only then changes which image it will try to boot. This separates ‘prepare the replacement’ from ‘make the replacement active,’ so a download or write failure need not destroy the last known-good path. [1][3]
The switching mechanism itself is security-critical. If an attacker can alter the activation record, corrupt both copies, or force repeated transitions, the presence of two slots is not enough. Real products may use A/B slots, a recovery partition, external service tools, or another design; each has storage, wear, complexity, and availability tradeoffs. [1][3]
How can A/B slots provide a way back?
In an A/B design, slot A contains the current working firmware while the candidate is written to slot B. The device first checks that the new image is complete and authorized. It then marks B as a trial boot rather than immediately discarding A. The exact metadata and state machine are product-specific and must themselves be protected. [1][3]
On the first trial boot, the device can perform a bounded health check: did essential services start, can the sensor read its input, and can the station report status? If the checks pass, the device records the new slot as accepted. If power fails or the trial never reaches an approved success condition, the boot policy may try the prior known-good slot or enter a designed recovery path. [1][3]
This is not a promise that every A/B device will always roll back automatically. The product must define when a trial counts as successful, how many attempts are allowed, what state survives reset, and whether either slot can be modified without authorization. A/B is one resilience pattern, not a substitute for signature checks or a complete recovery plan. [1][3]
An installation sequence and a minimum boot security version can be separate protected fields. This page's custom experiment commits the active slot and accepted sequence atomically after trial confirmation, retaining A and sequence 112 beforehand. That is one policy, not a universal ordering rule. Raising a real boot-version floor too early could also block A, the intended recovery image. Design the floors, trial state, permitted recovery versions and interrupted commit together.
What should a trial boot prove before it is accepted?
‘The processor started’ is not the same as ‘the product is healthy.’ A candidate image might reach a login prompt while its sensor driver is broken, its stored data cannot be read, or its network interface no longer works. For the water station, those are meaningful failures even if the bootloader reports success.
Choose a small set of product-specific checks that show the device can perform its essential job. For example, the station can initialize the sensor, take a plausible reading, preserve required configuration, and send a signed status message to its monitoring service. A watchdog can help detect a system that stops making progress, but a watchdog timeout alone does not establish that the measurement is correct. [1][3]
Keep the acceptance rule narrow and explicit. A trial should not be allowed to rewrite the trusted update policy or erase the old recovery route before the required checks pass. If a check fails, the safe behavior depends on the product: retry, return to the prior slot, remain in restricted service, or request operator recovery. Some failures will interrupt service, and that operational cost belongs in the design. [1]
If both the update and the network fail, what then?
A recovery route matters most when ordinary installation has already failed. The station may have lost power while writing, rejected a candidate at trial boot, or become unable to reach its update server. Recovery might use a preserved known-good image, a local service interface, or a newly downloaded authorized package when connectivity returns. [1][3]
Recovery is still a security-sensitive boot path. It must authenticate the code it runs and apply a defined policy to the image it installs; otherwise, an attacker could deliberately trigger recovery and use it as a way around normal verification. Any recovery credential, physical service procedure, or alternate update channel needs an owner, access limits, and a revocation or maintenance plan. [1][3]
A resilient product should also preserve useful failure evidence: which check failed, which release was attempted, and whether the device switched slots. Logs must not leak secrets or become an unauthenticated command channel. Recovery can restore a safe operating state, but it may require a technician and may interrupt service; the article's promise is a planned path, not zero downtime. [1][3]
Can an authorized exception allow older firmware?
Sometimes a newer release introduces a defect, or a site needs a previously qualified version while engineers investigate. A blanket rule that makes all rollback impossible can turn a release mistake into a prolonged outage. Yet simply turning off rollback checks invites an attacker to reinstall a vulnerable image.
A controlled exception can preserve the distinction between ‘old’ and ‘unauthorized.’ For example, an authorized publisher may issue a new manifest with a higher sequence number that explicitly permits a particular older image for a defined device class and recovery purpose. The device still checks the signer, scope, image digest, and current sequence policy; it does not accept any arbitrary old package. [2]
The operator should be able to explain who approved the exception, why it was needed, which devices and images it covers, and when it expires or is replaced. This creates an auditable decision instead of a hidden policy bypass. The exact sequence and exception model depends on the update architecture; RFC 9124 gives one manifest-based framework rather than a universal product rule. [2]
What if the signing key itself must change?
A firmware publisher may need to rotate a signing key because it is nearing retirement, its custody has changed, or compromise is suspected. Devices that trust only the old key need a secure way to learn which new key is authorized. A random key file delivered beside the firmware cannot establish its own authority.
A planned transition can use an already trusted authority to authorize the replacement key, then verify that the device received and stored the new trust information before depending on it. The update format may support key identifiers, sequence policy, and revocation information, but the device still needs a protected root or another independently trusted recovery route to validate that transition. [1][2]
If the root authority is compromised, merely signing a new key with that same root does not restore trust. The response may require a separately protected recovery mechanism, physical service, or a vendor-specific process. Teams should test routine rotation and emergency revocation before an incident, document who can approve each step, and retain evidence of what trust state was installed. [1][2]
Return to the remote water station
The station receives the security fix again. It verifies the authorized signer, checks that the manifest targets its model, confirms the digest of the staged image, and applies its protected sequence policy. The current firmware remains available while the candidate is written and checked; only after the designed activation step does the station attempt the new release.
During trial boot, the station checks the sensor and reporting path. If the release meets the acceptance rule, it records success. If installation or health checks fail, it follows the recovery behavior the product team designed—even if that means an operator visit or temporary loss of readings. A safe design makes these outcomes visible rather than pretending failures cannot happen.
The update path has several responsibilities: authenticate the authority, bind the image to the target, prevent disallowed replay, install without discarding the recovery path, test the product's essential function, and maintain a controlled route for recovery and key changes. Secure Boot can later enforce what may execute at startup, but secure updating must also protect the path that changes the firmware. [1][2][3]
Five points to keep
- A download server delivers a candidate; a device must independently verify the authorized signer and scope.
- A valid signature does not prove that an image targets this model or remains acceptable under current policy.
- A protected manifest sequence can reject stale replay; it is separate from the human-readable firmware version.
- Staging, A/B slots, and trial boot are design patterns that can preserve a recovery option, not guarantees of zero downtime.
- Recovery and signing-key changes are security paths too; they need authorization, verification, ownership, and testing.
Next, examine debug interfaces and physical access: which protections remain when an attacker can reach the device itself?
References
- IETF RFC 9019: A Firmware Update Architecture for Internet of Things: Update architecture, authorization, installation, failure handling, and recovery considerations.
- IETF RFC 9124, Section 4.3.1: Monotonic Sequence Numbers: Explains that the sequence is not a firmware version and that a later manifest sequence may authorize an older firmware version.
- NIST SP 800-193: Platform Firmware Resiliency Guidelines: Protect, detect, and recover principles for platform firmware resiliency.
Power fails during update: which slot can boot?
Start with active A:v7 and accepted sequence 112. Interrupt writing or trial boot, then run a confirmed update. Compare when the active slot and sequence floor commit together. Set sequence 112 and inspect immediate replay rejection.
This is one explicit A/B product policy with supplied signature and target results. Firmware v8 and sequence 113 are separate fields, not a universal version rule. Commit is modeled as atomic and protected. Real flash behavior, power-loss windows, trial watchdogs, recovery and floor ordering need separate design and validation. This is neither a SUIT parser nor proof of power-loss safety.
Learning guide
How Trust Is Built Inside a Chip
Open the course outline → · Progress counts published lessons only
Prerequisites
- Secure Boot and digital signatures are helpful background
What I learned
- Separate package delivery from authorization and compatibility
- Explain why valid signatures do not prevent replay of an old release
- Compare staged installation, trial boot, and recovery