< Academy

Why LLMs Struggle With Spreadsheets and How SpreadsheetLLM Fixes It

Research
Dr Nadine Kroher
Chief Scientific Officer

‍

‍

Most engineers, ourselves included, aren't big fans of spreadsheets. They're hard enough for humans to parse at a glance; they're even harder for machines. LLMs have historically struggled with them for the same reason. In a recent Passion Academy session, we looked at why that is and at SpreadsheetLLM, a paper from a group of researchers at Microsoft that shows how the problem can be overcome.

‍

What's Changed Since 2023?

‍

Back in 2023, when the first consumer-grade LLMs became available, dropping in an Excel sheet and asking a question rarely got you a reasonable answer. Today, even on fairly complex sheets with several sub-sheets and many tables, most LLMs do a decent job. The natural assumption is that the models simply got better. That's part of it, but not the whole story, and it's the other part that the session focused on.

‍

‍

Why Spreadsheets Are Hard for LLMs

‍

‍

The paper worked with a spreadsheet containing more than 60,000 tokens, which is already a large context on its own. But context size isn't actually the biggest issue. A few structural properties of spreadsheets make them difficult regardless of size:

‍

‍

  • They're two-dimensional: The position and coordinates of content carry real meaning but an LLM can only take in linear text. Feeding a spreadsheet to a model loses that positioning.
  • A single sheet often contains many tables: There's no guarantee of one clean table per sheet.
  • Formatting is used as data: A cell's colour or font can carry meaning in its own right, not just the value inside it.
  • Sheets are full of empty cells: People rarely trim the default empty rows and columns that appear when a spreadsheet is first opened.

‍

‍

All of this combines to make naive approaches fail. If you simply convert a spreadsheet to text, giving the LLM the row index, column index and cell contents for each cell, the results are poor. Interestingly, adding formatting information to that same linear text doesn't meaningfully improve things either. Feeding the data in linearly, even dressed up with style and formatting, just doesn't work well.

‍

‍

The Fix: Compress Before You Reason

‍

‍

SpreadsheetLLM's approach is to pass every spreadsheet through a compression stage before any reasoning happens. That stage has three modules, each creating a different compressed representation of the same data.

‍

‍

  1. Structural anchor detection. A heuristic algorithm, no LLM or AI involved, finds header rows and isolates them, then keeps only a handful of example rows. This is lossy compression: you don't retain everything, but you get a clear picture of where different tables sit within the sheet.‍
  2. Inverted-index value compression. Since spreadsheets have so many empty cells and repeated values, cell contents are encoded with the value as a key and the cell locations where it appears as the value. This is lossless compression, a technique familiar from sparse matrices elsewhere in engineering, and it reduces the data representation substantially on its own.
  3. ‍Data-type abstraction. The third representation drops the actual values entirely and instead records what type of data sits where: this block is integers, that block is dates. This is lossy again, but gives an extremely compact view of the sheet's structure.

‍

Together, these three representations compressed the context by a factor of 25 in the paper's experiments, while preserving the information actually needed to locate and answer questions.

‍

‍

How It's Used to Answer a Question

‍

‍

The process runs in stages. First, the compressed representations are computed for the whole sheet. Given a user query, the system uses the first (structural) representation to locate the relevant table. Then it uses the other representations to narrow down to the exact area, or areas, needed to answer the question. Only at this point does the LLM get involved, and only with that narrowed-down region, which can be chunked further if it's still large. The answer is then produced based only on the information relevant to the question, not the entire sheet.

‍

One technique that isn't in this paper but is common in similar systems: when a user asks for a calculation (such as an average or standard deviation, or wants a single value extracted) it often makes sense to have the LLM write code, or call pre-built code tools, rather than reason its way to the number directly. This makes the result deterministic, auditable and highly scalable, since a piece of code can handle tens of thousands of rows almost instantly, in a way that asking an LLM to do the arithmetic directly cannot match.

‍

‍

Does It Actually Work?

‍

The comparison that matters most is against a vanilla GPT-4 run on the raw sheet, which scored 52 on a micro-F1 metric for correctly identifying the relevant table. Adding the compressed representations produced a significant improvement on that table-detection task, and a major jump on the full question-answering task.

‍

The most striking result came from the 22 largest sheets in the benchmark dataset used in the paper. On those, naive GPT-4, fed a plain linearised text version of the sheet, answered none of the questions correctly. Using the compression approach answered a meaningful share of them.

‍

‍

The Real Takeaway

‍

‍

The important thing about this comparison is that the underlying model is exactly the same in both conditions. Nothing about GPT-4 itself changed between the naive approach and the compressed one. The entire improvement comes from what surrounds the model: compression, encoding, pre-computed representations and smarter tools.

‍

That has a broader implication worth sitting with. When a new version of your favourite LLM launches and it performs better, it's natural to assume the model itself improved. Often it did. But it's also very likely that the periphery around it got more powerful too: more tools it can call, more built-in functions, cleverer algorithms for handling specific types of data. Building intelligent periphery around a model can improve its practical performance as much as, or more than, waiting for the next model generation.

‍

‍

References

‍

  • Microsoft Research, SpreadsheetLLM (paper referenced in session)

‍

< back to academy
< previous
Next >