[PATCH] asn1_compiler: fix heap overflows on small or token-dense grammars

From: lzhan011

Date: Mon Oct 05 2026 - 05:01:06 EST


tokenise() allocates (len / 2) tokens on the assumption that there are
never more tokens than half the number of input characters. That doesn't
hold: single-character tokens may be adjacent, so "{}" yields two tokens
from two bytes while only one slot is allocated, and a 1-byte grammar
gets a zero-sized array. The tokeniser then writes past the end of the
heap buffer:

AddressSanitizer: heap-buffer-overflow
WRITE of size 2 ... in tokenise scripts/asn1_compiler.c:407

Every token consumes at least one character, so size the array by the
input length instead. Allocate one extra, zeroed slot as well, since
build_type_list() points the stop-marker type at token_list[nr_tokens].

Additionally, build_type_list() iterates with "n < nr_tokens - 1" where
nr_tokens is unsigned. A grammar that contains no tokens (for example a
file holding only a comment) makes this underflow and read far beyond
token_list:

AddressSanitizer: heap-buffer-overflow
READ of size 1 ... in build_type_list scripts/asn1_compiler.c:753

Use "n + 1 < nr_tokens" instead, so that such input is rejected with the
existing "No defined types" error.

The generated output for all in-tree .asn1 grammars is unchanged.

Found by fuzzing asn1_compiler built with ASan/UBSan.

Reproducers:
printf '{}' > a.asn1 && scripts/asn1_compiler a.asn1 a.c a.h
printf -- '-- x\n' > b.asn1 && scripts/asn1_compiler b.asn1 b.c b.h

Fixes: 4520c6a49af8 ("X.509: Add simple ASN.1 grammar compiler")
Assisted-by: Claude:claude-opus-5-5 ASan UBSan
Signed-off-by: lzhan011 <lzsx618@xxxxxxxxx>
---
scripts/asn1_compiler.c | 10 +++++-----
1 file changed, 5 insertions(+), 5 deletions(-)

diff --git a/scripts/asn1_compiler.c b/scripts/asn1_compiler.c
index 4c3f64506..1ae89637d 100644
--- a/scripts/asn1_compiler.c
+++ b/scripts/asn1_compiler.c
@@ -349,10 +349,10 @@ static void tokenise(char *buffer, char *end)
char *line, *nl, *start, *p, *q;
unsigned tix, lineno;

- /* Assume we're going to have half as many tokens as we have
- * characters
+ /* Every token consumes at least one character, so we can't have
+ * more tokens than we have characters
*/
- token_list = tokens = calloc((end - buffer) / 2, sizeof(struct token));
+ token_list = tokens = calloc(end - buffer + 1, sizeof(struct token));
if (!tokens) {
perror(NULL);
exit(1);
@@ -749,7 +749,7 @@ static void build_type_list(void)
unsigned nr, t, n;

nr = 0;
- for (n = 0; n < nr_tokens - 1; n++)
+ for (n = 0; n + 1 < nr_tokens; n++)
if (token_list[n + 0].token_type == TOKEN_TYPE_NAME &&
token_list[n + 1].token_type == TOKEN_ASSIGNMENT)
nr++;
@@ -773,7 +773,7 @@ static void build_type_list(void)

t = 0;
types[t].flags |= TYPE_BEGIN;
- for (n = 0; n < nr_tokens - 1; n++) {
+ for (n = 0; n + 1 < nr_tokens; n++) {
if (token_list[n + 0].token_type == TOKEN_TYPE_NAME &&
token_list[n + 1].token_type == TOKEN_ASSIGNMENT) {
types[t].name = &token_list[n];
--
2.34.1