
A critical Light Protocol initialization bug, hidden Anchor instructions, and a blind spot in AI audits.
Last year we audited an upgrade to Light Protocol that was moving its compressed token program from Anchor to Pinocchio. During that audit, we found a critical bug with a fairly unusual setup, and we'd like to walk you through it today.
There are a couple of reasons we want to talk about this one now. With hackhack.ai, we're building AI tooling to find bugs, and in our benchmarking backtests, AI auditors keep missing this issue. We'd like to understand why.
Second, this bug gets particularly interesting when you look at the program upgrade itself. We'll get to that, and the AI part, after walking through how it works.
Light has since fixed the issue. But let's dive into how Light Protocol worked at the time of our audit, and how an account meant to store an IDL could end up being treated as a token account.
Light Protocol builds a compression layer on Solana. Basically, it allows state, such as token accounts, to be represented in Merkle trees instead of storing every record in its own Solana account. On chain, Light keeps the tree roots and the supporting tree and queue state needed to validate updates. That reduces the storage and rent costs of maintaining lots of individual accounts.
The full records remain available through the ledger and indexers. When you use a compressed account, you supply its data and the evidence needed to authenticate it against the on-chain commitments. A transaction consumes the old state and creates replacement state, with checks that prevent the old state from being spent again. Light's protocol paper explains how those pieces fit together.
Light's compressed token program, which we'll call the cToken program, applies that model to tokens. It implements operations such as minting, transferring, and burning compressed tokens. Moving tokens between holders still requires authority and balance checks; compression changes how the token state is stored and authenticated. You can see those operations in the older Anchor program.
So why would this program also need ordinary token accounts? An ordinary account gives a token balance an on-chain address and a data record that programs can read and update directly. Light supported that form too, using a layout compatible with SPL token accounts. Apps could use familiar account-based operations, including direct transfers, while balances could move between ordinary and compressed form through compression and decompression.
A token account answers a few basic questions: which token does it hold, who may spend it, and how many units does it hold? The token is identified by its mint. Its spending authority is usually called its owner.
That second name is easy to confuse with Solana account ownership. The Solana account's owner is the program allowed to modify its data. The token record's owner is an authority the token program recognizes inside that data. A wallet can control a token balance while the account storing that balance belongs to a token program. Solana's account documentation describes the runtime ownership rule.
At the time of our audit, Light could represent a token balance in two ways:
The balance can move between these two forms. Creating an empty account shouldn't create any token value.
In the code, these are CToken and TokenData. When we say “cToken account” below, we mean the ordinary Solana account. This is the form whose initializer contained the bug.
For the cToken path, compression reduces an ordinary account's balance while creating corresponding compressed value. Decompression consumes compressed value and increases the ordinary balance. The historical compression and decompression code implements the ordinary-account side of those operations.
Moving between these representations must preserve value. It must not manufacture a balance. The same applies when creating an empty account.
When you create a cToken account, initialize_ctoken_account is supposed to fill in its mint and spending authority and give it a zero balance with default settings. That's also what its documentation says. To issue tokens, you use a separate mint-to-cToken handler, which checks the mint authority and updates supply.
First, though, the account needs storage. Allocating that storage and initializing the token fields are separate steps, and they don't have to happen in the same instruction.
The creation handler supported an ordinary form, compatible with SPL InitializeAccount3, which accepted already allocated storage. It also supported an extended form with configuration for Light's compression and rent lifecycle. The ordinary form is enough to understand this finding.
A normal client could allocate a fresh account, then initialize it using Light's SDK. This is a shortened example using the historical CreateCTokenAccount builder; imports, rent calculation, and transaction setup are omitted:
let allocate = system_instruction::create_account(
&payer,
&new_account,
rent_exempt_balance,
base_account_size as u64,
&ctoken_program_id,
);
let initialize = CreateCTokenAccount {
payer,
account: new_account,
mint,
owner: token_authority,
compressible: None,
}
.instruction()?;
// Submit [allocate, initialize] in one transaction.
// Sign with the payer's and new account's keypairs.
base_account_size is the size of the ordinary token record, and ctoken_program_id is Light's compressed token program. compressible: None leaves out the optional extension. Fresh allocation gives this flow zeroed storage. The historical compatibility test also demonstrates allocation followed by initialization.
The base token record follows the SPL account layout. It contains the mint, spending authority, amount, delegate information, account state, native-token reserve information, and close authority. Unlike an Anchor account, it does not begin with an eight-byte Anchor type discriminator. The historical type definition describes these fields.
The vulnerable shared initializer checked that the account had the expected length and that its token-state marker said uninitialized. Its writes to the base token record amounted to this:
// Schematic: the implementation writes directly into the byte buffer.
token.mint = mint;
token.owner = owner;
token.state = Initialized;
It left the amount and other base fields alone.
The comment above those writes captured the assumption:
Account is already zeroed, only need to set these 3 fields
That would be fine if every account it accepted were freshly allocated, untouched storage. But the initializer did not check that.
An uninitialized marker is a field value. It is not evidence that the entire account is empty.
The problem is that this program owns more than just token accounts. Knowing that an account belongs to the program doesn't tell us what's already inside it, or which instructions put it there.
Our earlier post, Hidden IDL Instructions and How To Abuse Them, explains why the instructions written in an Anchor program's main source are not necessarily its complete interface. Anchor can generate additional instructions for storing and updating an on-chain IDL.
An IDL describes a program's interface: its instructions, arguments, and account types. An on-chain copy lets tooling discover that interface from a program address. Anchor also supports temporary IDL buffers, so a large IDL can be uploaded in chunks before replacing the canonical copy.
Light was using Anchor 0.31.1 at the time. Let's look at how that version handled IDL accounts and buffers.
The canonical IDL account has a deterministic address derived using the program and the anchor:idl seed. A temporary buffer is a separate account. Creating a buffer does not mean changing the canonical IDL, and controlling a buffer does not require control of the program's upgrade authority.
The generated buffer handler records a signer as the buffer authority. Subsequent writes require that authority. Copying a buffer into the canonical IDL has additional authority checks. Those are different permissions: a user can control a buffer without being allowed to replace the project's published IDL. Anchor's generated handlers show these checks.
In that version, the IDL account has the following format. Offsets are relative to the account's data, and ranges include both endpoints.
The program owns the account, but the buffer authority can write its payload. The recorded length can be smaller than the space allocated for it.
The fixed header is 44 bytes. The payload is trailing storage rather than an explicitly declared Vec field in the Rust account struct. Anchor's generated accessors take the payload slice after that header. Tooling uses this space for a compressed IDL file; that file compression is separate from Light's state compression. Source: generated IdlAccount and trailing-data accessors.
data_len tracks uploaded payload length. It is not the total account length, and it does not have to equal the allocated payload capacity. The write handler appends data and advances that length.
The buffer's Solana owner is the application program. But a user who controls the buffer authority can influence its payload. The program's own generated handlers perform the writes, so this follows Solana's ownership rules.
So “owned by our program” and “contains user-influenced data” can both be true.
Now we can put the two pieces together. Anchor lets a user control the contents of a buffer owned by the program. The token initializer sees an account of the expected size with an uninitialized token-state marker, and fills in only some of its fields.
The account still carries its IDL discriminator. The token initializer doesn't check it because it expects an SPL-compatible layout, which has no Anchor discriminator. It could therefore treat an existing account of another type as storage awaiting token initialization.
Successful initialization had to establish a zero balance. Instead, a pre-existing value could survive and become the newly initialized account's amount. The account could end up with a caller-selected mint and spending authority while retaining a balance that was never credited through an authorized token operation. This is the impact documented in ACC-C1.
The blue fields get written. The pink fields keep their existing values. A marker saying “uninitialized” doesn't tell us those values are zero.
We're taking an account that was already initialized as one type and initializing it again as another. This combines type confusion with incomplete initialization. We don't need to reset somebody else's token account first. The account only has to look uninitialized when the token initializer reads it.
The result is a token balance that nobody minted or deposited. Once other operations accept that balance, their accounting is already wrong. For assets with redemption or exchange paths, that can expose backing assets or counterparties, depending on the checks and liquidity on those paths. This was an audit finding; we're not describing an observed theft or putting a dollar figure on the loss.
Compression cannot repair that accounting error. It authenticates state according to the checks it performs; it does not independently establish that every earlier program operation created the value legitimately. The token-account initializer needed to establish a valid starting balance.
The IDL-buffer machinery was already there in the Anchor program. An IDL buffer on its own doesn't give anyone a token balance. Something has to accept its bytes as token state before that becomes possible.
The Pinocchio migration added the ordinary cToken creation handler and its shared initializer. The Anchor predecessor had compressed-token operations and SPL token pools, but no corresponding ordinary-cToken creation entrypoint. The upgrade introduced the code that could give an existing buffer this new meaning.
That explains why the old buffer feature became dangerous during this change. The new initializer manually wrote a few fields and assumed the rest were zero. Its input could include accounts created for a completely different purpose. The failure was in that assumption; the same mistake could be made in a custom Anchor handler too.
Now consider a successor version that removes Anchor's IDL handlers entirely.
A review of that version might ask whether any of its instructions can create a program-owned IDL buffer. The answer could be no. But that would not establish that no such account can be passed to it.
Solana stores mutable application state in data accounts separate from executable program code. Replacing the executable at the same program address does not, by itself, enumerate and reset those accounts. Their ownership can still name the same program. Solana's program model explains this separation.
The relevant lifecycle is:
The code changes at the upgrade. The account doesn't disappear with the instruction that created it.
In that scenario, an adversary can prepare state before the upgrade and wait for the successor to interpret it unsafely. No IDL-buffer instruction needs to remain available in the successor.
The security requirement is broader than “the new program cannot create this state.” The new program must safely handle the states that earlier versions could leave behind, including states created through optional or generated framework features.
There's a wrinkle in the version we audited: Light hadn't completely removed Anchor yet. Its entrypoint still forwarded some instructions to Anchor, so the IDL machinery was relevant within that version too. Our audit report covered both possibilities: keeping those handlers and removing them.
The prepare-before-the-upgrade scenario matters when those handlers are removed. That's the case we're exploring here, rather than claiming that a deployed version followed that exact sequence. The old accounts can survive either way. Upgrades carry other version-transition traps too; we cover the toolchain and rollback side in Move your Solana program to sBPF v3 while you have time.
The linked remediation commit rejects account data containing any nonzero byte before initialization. That makes the assumption explicit: the initializer accepts empty storage, rather than trusting one field to stand in for the entire record.
By the time of that patch, the code already cleared the base token bytes before initializing them. The patch changed this to check that the data was already zero and reject it otherwise. So an unexpected account gets rejected instead of having its contents silently overwritten.
The same commit also tightened up account creation: if an account was already funded, it had to belong to the System Program before the program could assign and resize it. That's a separate check from making sure the initializer receives zeroed data.
For an account format without a type discriminator, the initializer has to check that initialization is allowed, then set every field to its expected starting value. Clearing bytes is only safe when the account is eligible for initialization. Otherwise, it could destroy a valid account of another type.
Removing IDL support can reduce the ways future accounts are created. It cannot substitute for validating inherited accounts.
This is the part we've been thinking about while building our AI auditing tools. We keep seeing this finding missed in our backtests. We have a suspicion about why, but we don't have a proven explanation.
An auditor looking at the initializer can see that it leaves some fields untouched. The next question is whether those fields can contain anything other than zero. If the audit only follows the account-creation paths in the new code, it might conclude that they can't. Every account it has considered starts as fresh storage.
But the program doesn't get a fresh set of accounts when it's upgraded. It inherits whatever the old version left behind. In the version without IDL handlers, the relevant account-creation path has disappeared from the code being reviewed. To understand which inputs the new initializer can receive, you have to consider behavior that belonged to a different version of the program.
That's our main suspicion: the AI may be reasoning about the new code with too narrow an idea of which accounts can exist. The interesting part is the combination of what the old program could create and what the new program would accept.
There are other possible explanations. It could miss Anchor's generated IDL instructions entirely, or see them without connecting the buffer format to the token initializer. It could also trust the comment saying the account was already zeroed. Having the framework source available doesn't guarantee that an auditor follows every generated instruction and connects it to the application code. And because the version we audited still had an Anchor fallback, missing those handlers alone could explain a miss there.
After the upgrade, the problematic account can already exist and be exploitable. What the new code alone may not explain is how it got there.
This is one of the things we want AI auditors to get better at. When an audit includes a program upgrade, it needs to reason about the state the program inherits, including accounts created by framework features that may no longer exist. Reading the current implementation carefully is only part of that job.
For this bug, the useful question was simple: the comment says the account is already zeroed, but what guarantees that? Follow that question through the old code, the new code, and the generated framework instructions, and the assumption falls apart.