conrads.website

Parsing LLM Output Streams

Notes from guiding a coding agent through parsing structured output as it streams.

Code: pocket_parsing

Introduction

Motivation

Can we parse the code output of an LLM agent while it's streaming, to validate correctness closer to real time?

Can we mitigate this behavior by counting errors over time?

For example, when an erroneous error like this is detected, it is likely to be resolved soon. When resolved, the error count will go back to 0. Perhaps an error threshold will work, because when errors arise, if they are erroneous they will disappear before many errors accumulate.

Could lark work for incremental parsing?

Results

Tree-sitter reports errors on every half-written line, so an error count over time is too noisy to stop a stream on: the plot shows spikes through every multi-line block. Lark was more promising. Its exception lists the tokens it expected next, which separates “the input just ended” from “this is wrong”, and with extra newlines in the examples it behaved as hoped. It is still a toy. The experiment is in pocket_parsing.


There's some useful potential for wrapping parsers directly around streamed outputs for better evaluation. Usually, Cline just makes a best attempt and then fails, executes the bad code, and parses the errors. This is not the direction I take as a programmer.

Documentation access is also usually ignored as a step, both by Cline and the model itself (even when documentation is explicitly provided).