Insight

Why We Chose a 28-Byte Backup Header

A backup format is more than ciphertext. Signet's 28-byte header makes each backup self-describing, upgradeable, and tamper-evident without hiding the parameters needed to derive its key.

Attomus insight

A backup format is more than ciphertext. Signet's 28-byte header makes each backup self-describing, upgradeable, and tamper-evident without hiding the parameters needed to derive its key.

Part 3 of 5 in Building Signet

The backup format for an authenticator app looks like a solved problem. Encrypt the secrets with a passphrase, write the ciphertext to a file, done.

It is not done. The naive version fails silently in several distinct ways, and the worst of them leaves a user holding a file that will never decrypt again, reporting an error that blames their passphrase for something else entirely. The 28-byte header we settled on came out of working through those failures one at a time. The same reasoning applies to any format that combines key derivation with authenticated encryption.

What the naive format looks like

The first version of any passphrase-based backup format, with a random salt, looks roughly like this:

[salt: 16 bytes][nonce: 12 bytes][ciphertext + tag: n + 16 bytes]

Derive a key from the passphrase and the salt using whichever KDF you chose, encrypt with AES-256-GCM, write it out. Nothing there is wrong, exactly. The problem is what it leaves unsaid.

The KDF algorithm, the KDF parameters and the cipher are all implicit, hardcoded somewhere in the application. Change any of them – stronger Argon2 parameters for new exports, say – and you have a versioning problem with no version information in the file to solve it with. The fix is a lookup table in your code: format version 1 used these parameters, version 2 used those. That table has to live forever. Deprecate an old version of the app and you have quietly deprecated the ability to read the backups it wrote, with nothing in the file to say so.

The failure mode is quiet, and it arrives years later. The user has a backup file. They reinstall the app two years after writing it. They enter the correct passphrase. The decrypt fails, and the message says “incorrect passphrase” – because that is the only failure the code knows how to report. The file is from a format version the new app no longer understands, and nothing in it can tell anyone that.

Adding a version byte makes it worse

The obvious fix is a version byte:

[version: 1 byte][salt: 16 bytes][nonce: 12 bytes][ciphertext + tag: n + 16 bytes]

The app reads the version and dispatches to the right decryption path, which solves the immediate problem and introduces a new one. The version byte is unauthenticated. It sits outside the AEAD ciphertext, so an attacker who can modify the file can flip it and push the code down a different decryption path.

The practical damage is limited: the key derived from the wrong parameters will not authenticate the ciphertext, so decryption fails. But “an attacker can make decryption fail” is still an unwanted property in a file users are trusting to be tamper-evident, and the principle it breaks is a general one. Authentication should cover the whole file, including every field that controls how the file is processed.

The version byte is also a proxy for the underlying problem rather than a solution to it. The underlying problem is that the KDF parameters are not in the file.

KDF parameters belong in the file, not in the code

A backup file has to be self-describing for the purpose of decryption: anything required to decrypt it must be carried in it. That means the KDF algorithm, the time cost, the memory cost, the parallelism factor and the salt. If any of those are implicit, the file is not a backup. It is a backup plus a dependency on a particular version of your application.

The header we arrived at:

Offset  Size   Field
----------------------------------------------------------------
0       4      magic bytes: 0x41 0x55 0x54 0x48 ("AUTH")
4       1      format version: 0x01
5       1      kdf identifier: 0x01 = Argon2id
6       1      argon2 time cost (1..255)
7       1      argon2 memory exponent (memory = 2^n KiB)
8       1      argon2 parallelism
9..24   16     salt (random per export)
25      1      cipher identifier: 0x01 = AES-256-GCM
26..27  2      reserved (zeroed; authenticated)
----------------------------------------------------------------
Total:  28 bytes

Followed by a 12-byte nonce, then the GCM ciphertext and its 16-byte authentication tag.

The 28-byte Signet backup header remains readable while AES-GCM authenticates every field as additional authenticated data.

The memory exponent encodes Argon2’s memory cost as a power of two, which lets one byte cover everything from 2 KiB at exponent 1 to 128 GiB at exponent 27 without wasting space on a value that will only ever hold round numbers. The current production value is 17 – 128 MiB – which raises the cost of GPU-parallel guessing materially whilst still completing in under a second on a mid-range phone.

The header must be part of the authenticated data

The piece that makes the format properly tamper-evident is that the whole 28-byte header is passed as additional authenticated data (AAD) to the AES-GCM encryption. None of the header is encrypted, because the KDF parameters must be readable before the key can be derived. All of it is authenticated.

AES-GCM computes its 16-byte tag over the ciphertext and any AAD supplied, so any modification to the header produces a tag verification failure. The version byte cannot be flipped, the KDF parameters cannot be weakened, and the cipher identifier cannot be altered without breaking authentication.

fun encode(plaintext: ByteArray, passphrase: CharArray, params: ArgonParams): ByteArray {
    val salt = secureRandom(16)
    val header = buildHeader(salt, params)            // 28 bytes
    val keyBytes = argon2id(passphrase, salt, params) // derived from header fields
    val key = SecretKeySpec(keyBytes, "AES")
    val nonce = secureRandom(12)
    val cipher = Cipher.getInstance("AES/GCM/NoPadding")
    cipher.init(Cipher.ENCRYPT_MODE, key, GCMParameterSpec(128, nonce))
    cipher.updateAAD(header)                          // authenticated, not encrypted
    val ciphertext = cipher.doFinal(plaintext)        // includes the GCM tag
    return header + nonce + ciphertext
}

fun decode(data: ByteArray, passphrase: CharArray): ByteArray {
    val header = data.copyOfRange(0, 28)
    val params = parseHeader(header)
    val salt = header.copyOfRange(9, 25)
    val keyBytes = argon2id(passphrase, salt, params)
    val key = SecretKeySpec(keyBytes, "AES")
    val nonce = data.copyOfRange(28, 40)
    val ciphertext = data.copyOfRange(40, data.size)
    val cipher = Cipher.getInstance("AES/GCM/NoPadding")
    cipher.init(Cipher.DECRYPT_MODE, key, GCMParameterSpec(128, nonce))
    cipher.updateAAD(header)
    return cipher.doFinal(ciphertext)                 // fails on any tampering
}

The magic bytes at the front earn their place on practical grounds rather than security ones. They allow the implementation to fail fast on a file that was never a backup – a photo, a text document, an export in the old format – before it starts key derivation, which is deliberately expensive. Four bytes spent on a magic check saves several hundred milliseconds and returns a clear “wrong file type” instead of a misleading “incorrect passphrase”.

What the format buys you

A backup file in this format is self-describing against the published specification. A decryption implementation does not need to know which version of the app wrote the file: it reads the header, identifies the KDF algorithm and its parameters, derives the key from the passphrase, then authenticates and decrypts. Because the parameters travel with the file, they can be strengthened in later versions – new exports use the stronger settings, old exports still decrypt correctly from the header they carry. No lookup table, no version dependency.

The AAD binding makes the file tamper-evident end to end. The ciphertext, the nonce and the header are all covered by the authentication tag, and a file modified in any authenticated field fails authentication before a single byte of plaintext is returned.

The reserved bytes at the end are authenticated but currently zeroed. Version 1 implementations should reject a file whose reserved bytes are non-zero, rather than guessing at what they might mean. Every binary format that reaches production eventually has to evolve, and two spare bytes cost nothing now whilst being impossible to retrofit later, once files in the old shape are already in circulation. Rejecting unknown values keeps the room usable: a future version can define those fields deliberately, with the version byte making the change explicit.

Twenty-eight bytes of overhead, and the file explains itself to anything that can read the specification.

Attomus Signet — offline 2FA, on the App Store and Google Play. Read the Signet assurance record.