Why LLMs Struggle With Spreadsheets and How SpreadsheetLLM Fixes It


Most engineers, ourselves included, aren't big fans of spreadsheets. They're hard enough for humans to parse at a glance; they're even harder for machines. LLMs have historically struggled with them for the same reason. In a recent Passion Academy session, we looked at why that is and at SpreadsheetLLM, a paper from a group of researchers at Microsoft that shows how the problem can be overcome.
Back in 2023, when the first consumer-grade LLMs became available, dropping in an Excel sheet and asking a question rarely got you a reasonable answer. Today, even on fairly complex sheets with several sub-sheets and many tables, most LLMs do a decent job. The natural assumption is that the models simply got better. That's part of it, but not the whole story, and it's the other part that the session focused on.
The paper worked with a spreadsheet containing more than 60,000 tokens, which is already a large context on its own. But context size isn't actually the biggest issue. A few structural properties of spreadsheets make them difficult regardless of size:
All of this combines to make naive approaches fail. If you simply convert a spreadsheet to text, giving the LLM the row index, column index and cell contents for each cell, the results are poor. Interestingly, adding formatting information to that same linear text doesn't meaningfully improve things either. Feeding the data in linearly, even dressed up with style and formatting, just doesn't work well.
SpreadsheetLLM's approach is to pass every spreadsheet through a compression stage before any reasoning happens. That stage has three modules, each creating a different compressed representation of the same data.
Together, these three representations compressed the context by a factor of 25 in the paper's experiments, while preserving the information actually needed to locate and answer questions.
The process runs in stages. First, the compressed representations are computed for the whole sheet. Given a user query, the system uses the first (structural) representation to locate the relevant table. Then it uses the other representations to narrow down to the exact area, or areas, needed to answer the question. Only at this point does the LLM get involved, and only with that narrowed-down region, which can be chunked further if it's still large. The answer is then produced based only on the information relevant to the question, not the entire sheet.
One technique that isn't in this paper but is common in similar systems: when a user asks for a calculation (such as an average or standard deviation, or wants a single value extracted) it often makes sense to have the LLM write code, or call pre-built code tools, rather than reason its way to the number directly. This makes the result deterministic, auditable and highly scalable, since a piece of code can handle tens of thousands of rows almost instantly, in a way that asking an LLM to do the arithmetic directly cannot match.
The comparison that matters most is against a vanilla GPT-4 run on the raw sheet, which scored 52 on a micro-F1 metric for correctly identifying the relevant table. Adding the compressed representations produced a significant improvement on that table-detection task, and a major jump on the full question-answering task.
The most striking result came from the 22 largest sheets in the benchmark dataset used in the paper. On those, naive GPT-4, fed a plain linearised text version of the sheet, answered none of the questions correctly. Using the compression approach answered a meaningful share of them.
The important thing about this comparison is that the underlying model is exactly the same in both conditions. Nothing about GPT-4 itself changed between the naive approach and the compressed one. The entire improvement comes from what surrounds the model: compression, encoding, pre-computed representations and smarter tools.
That has a broader implication worth sitting with. When a new version of your favourite LLM launches and it performs better, it's natural to assume the model itself improved. Often it did. But it's also very likely that the periphery around it got more powerful too: more tools it can call, more built-in functions, cleverer algorithms for handling specific types of data. Building intelligent periphery around a model can improve its practical performance as much as, or more than, waiting for the next model generation.