8051: Building a Custom Disassembler
Industrializing the disassembly of an undocumented processor from a raw binary is a complex task that can be broken down into four key steps: Verify that the binary does not belong to a known processor. Verify that the binary is not obfuscated, compressed, or encrypted code for a known processor. Build an undocumented processor generator. Create the analysis pipeline and custom disassembler…
The story revolves around building a custom disassembler for an undocumented processor by breaking down the task into four key steps: verifying the binary, building the processor generator, creating the analysis pipeline, and developing the custom disassembler generation process. The focus is on creating lightweight disassemblers for bare-metal binaries, as tools like Ghidra require manual processor target selection before analysis can begin.
The article outlines three approaches to building a custom disassembler. The first approach, using Ghidra and SLAgh specification language, proved unsuitable due to metadata loss and lack of standardization in .slaspec files. The second approach, using dis51, also failed as it is an execution-tracing disassembler that requires valid program flow, not a simple linear sweep of opcodes.
The third and successful approach involved using an LLM (Gemini) to generate a normalized instruction mapping table. The model produced a Python script containing the opcode mapping and a processor-specific lookup table for Special Function Registers (SFR). Using this table, the static disassembler was written in under an hour, and the foundation for a dynamic disassembly simulator was laid.
Although some bugs were present in the generated table, the time saved was significant, and the core structure of the disassembler remained consistent when porting to other architectures.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.