Cleaning up a messy Excel spreadsheet can easily take longer than building it in the first place. I wondered whether Copilot could take care of the tedious work for me, so I gave it a deliberately chaotic workbook and a simple prompt. But instead of simply fixing problems, much of its effort went into deciding which ones it shouldn't touch.
Rather than spending ages crafting the perfect AI prompt, I wanted to see how far Microsoft's AI had really come. After all, Copilot is integrated directly into Excel. Unlike standalone AI tools, it already has access to the workbook, so I wanted to see how much it could understand without me explaining where the data lives or describing every column.
There's another reason I kept the prompt brief. One of the biggest criticisms of AI assistants is that you can spend so long refining prompts that you could've finished the job yourself. If cleaning a spreadsheet requires multiple rounds of prompting, clarifying, and correcting, much of the promised time savings disappears.
Clean up the spreadsheet while preserving the underlying data. Improve the formatting, fix inconsistencies where appropriate, identify anything that needs manual review, and explain every significant change you make. Don't make assumptions where the correct value is unclear.
The workbook contained hundreds of rows of deliberately messy data, including inconsistent formatting, spelling variations, impossible dates, random fonts and colors, blank rows, and several suspicious values. Microsoft recommends using structured tables for the best results, but I wanted to see how Copilot handled a genuinely messy workbook.
After analyzing the workbook, checking dates, looking for calculation problems, and verifying ambiguous values before making changes, Copilot declared the cleanup "complete."
The new "Issues – Manual Review" worksheet was the highlight of the test. I'd asked it to "identify anything that needs manual review," so I expected a brief summary. Instead, Copilot created an entire worksheet documenting problems needing human attention and grouped them into sensible categories.
More importantly, it explained why each item had been flagged. For example, impossible dates included explanations like "Month 13 does not exist" or "April only has 30 days." Calculation mismatches showed both the recorded total and the value Copilot expected based on the quantity and unit price. And for ambiguous dates, it offered possible interpretations, not blind guesses.
That kind of detail would save time during a manual review. Despite the vague instruction, it created a surprisingly thorough audit of the workbook and explained its reasoning when it chose not to make changes.
The first thing I thought when I saw the "cleaned" spreadsheet was how much still looked unchanged. Yes, Freeze Panes had been applied, whitespace had been removed from hundreds of cells, and number formats had been improved. But many of the most obvious formatting problems remained.
Random fonts and font sizes were still scattered throughout the workbook, and existing cell fill colors were untouched, making Copilot's own conditional formatting much harder to interpret.
It also claimed to have standardized date formatting, but many were still inconsistent. This was one of the areas where I expected more, especially because formatting cleanup is relatively straightforward. While I intentionally didn't give Copilot detailed formatting instructions, I expected it to recognize obvious inconsistencies.
Currency formatting improved for numeric values, but cells containing text such as "USD 12.00" were simply flagged instead of converted. Even relatively minor decisions, like centering some columns while left-aligning others, felt arbitrary rather than genuinely improving readability.
In several places, Copilot also seemed to prioritize the wrong jobs. The header row already looked perfectly presentable, yet it spent time redesigning it while leaving much messier formatting elsewhere in the workbook untouched.
One thing Copilot consistently did well was resisting the temptation to make risky assumptions. It deliberately refused to standardize status values like "Complete," "Completed," and "Done," explaining that they might represent different business states. It also left mixed order ID formats alone because both appeared valid.