Urgent.News

What's breaking now, across thousands of outlets.

Tech

chatstyle: a C++ core and a Python app for comparing writing styles in chats (Russian only, so far)

Disclosure. Most of the code, the tests and the first draft of this article were produced with an AI assistant. I set the goals, supplied the data (my own Russian Telegram chats), told it which accounts belong to the same person, and ran the app myself on a real case: comparing a friend's new account against five candidates. An early version ranked the true author third; after more methods were…

Chatstyle is a C++ core backed by a Python application designed to compare writing styles in chats. It is available in Russian only at the moment. The tool aims to determine which candidate's writing style more closely matches an unknown author's messages, providing a similarity estimate rather than definitive proof of authorship.

The numeric core of chatstyle is built in C++17, handling n-grams, Delta, Impostors, character language models, writing rhythm, and ensemble methods. These calculations are exposed to Python via pybind11 for application layer functionality such as preprocessing, the ensemble, Telegram export parsing, encrypted chat storage, reports, CLI, desktop UI, and evaluation scripts.

Chatstyle operates independently on your machine, requiring no servers or telemetry. Texts are stored encrypted using AES-256-GCM, but Telegram message cache encryption is not currently implemented. The core is compatible with Unicode, handling Russian characters and supporting Chinese and Japanese per character.

The tool was developed using real Russian Telegram chats, with no real names or texts exposed in the repository. Various tests were conducted, including tasks with known answers by comparing the same person's messages across different chats and time splits. The metric measured whether the true author ranked first in the comparison.

Results showed that the ensemble method with lexicon mode achieved the highest accuracy at 95%, while the style-only ensemble reached 91%. Burrows Delta scored 72%, General Impostors 58%, and TF-IDF cosine 56%. The effect of the number of candidates showed that the style-only mode weakened as the candidate pool grew, while the lexicon mode remained stable. Text length also impacted accuracy, with higher accuracy for longer texts.

A real-world case demonstrated that the initial version (using only cosine, Delta, and Impostors) incorrectly ranked the true author third out of five candidates. After incorporating additional methods and the ensemble, the tool correctly identified the true author first.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in Tech

We are no longer continuing the interview.

So late last night, I got an update from Wasmer, that they will no longer be continuing with my application. A bit sad, ngl, especially given friday wasnt by choice that I had to cancel the interview.

Why We’re Building Animal Watch 365: Solving the 24/7 Rescue Bottleneck with Tech

Hey everyone! 👋 As developers, we often build apps for productivity, e-commerce, or SaaS. But every once in a while, it's refreshing to tackle a problem that has a direct, tangible impact on the…

  • Animal Watch 365 addresses 24/7 animal rescue bottleneck.
  • System uses real-time alerts and constant monitoring.
  • Technology focuses on precise location data and urgency levels.

More from Monday 5 October →