PDF, ZIP and binary files streaming over Kafka with names intact.
PDF, ZIP and binary files streaming over Kafka with names intact. Day 07 of the WKafka Open-Source Engineering Series. Don't provision a cloud bucket when all you need is sending small PDFs and ZIPs down your pipeline. WKafka supports format='file' out of the box. The Pain Points We Faced Building S3/MinIO upload/download pipelines just to move 500KB report files Lost original filenames and…
A new open-source library called WKafka enables the seamless streaming of PDF, ZIP, and binary files over Kafka, while preserving the original file names intact. This innovative solution is part of the WKafka Open-Source Engineering Series, which kicked off on Day 07.
Traditional approaches involve provisioning cloud buckets solely to send small PDFs and ZIP files down the pipeline. However, WKafka offers a more efficient alternative by natively supporting the file format. This means that no additional steps are required to move 500KB report files, as the library handles the entire process effortlessly.
During the development of S3/MinIO upload/download pipelines, the original filenames and MIME types were lost when converting files to raw bytes. Moreover, the latency caused by polling for cloud object storage propagation became a significant issue. WKafka addresses these pain points by implementing a streamlined file serialization process that bundles the filename and binary data within the native file serializer.
To implement this solution, the producer side uses the `kafka.produce_file()` function, specifying the topic and file path. For instance, `kafka.produce_file(topic='reports', file_path='audit_2026.pdf')` sends the file 'audit_2026.pdf' to the 'reports' topic. On the consumer side, the `@kafka.consumer` decorator is used with the `format='file'` parameter.
Once a message is received, the `msg.save_file()` method is called to write the incoming file to disk, automatically preserving the original file name. For example, `msg.save_file(target_dir='received_pipeline_files/')` saves the file in the 'received_pipeline_files/' directory.
The architecture of WKafka provides several advantages. Firstly, the native file serializer combines the filename and binary data, eliminating the need for external storage solutions. Secondly, the `msg.save_file()` method allows for a one-liner to write the incoming file to disk, preserving its name. This streamlined process leads to sub-second delivery, eliminating bottlenecks that may arise from external storage.
The library has been rigorously tested and verified against real broker clusters, ensuring its reliability and compatibility.
WKafka is compatible with Python versions 3.9 through 3.14, maintaining strict typing throughout. For more information, developers can visit the official GitHub repository at https://github.com/wisrovi/wkafka or the project's PyPI page at https://pypi.org/project/wkafka. The library falls under the realms of #Python, #DataEngineering, #OpenSource, and proudly supports #Wisrovi.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.