Urgent.News

What's breaking now, across thousands of outlets.

AI

Show HN: I trained a 125M model to autocomplete piano on-device

I trained a 125M-parameter transformer to autocomplete piano performances in real time (~108 notes/sec on an iPhone 15). The idea is basically GitHub Copilot or Tabnine, except instead of prompting it with code, you prompt it by playing a few notes on a MIDI piano. The model then continues what you played, entirely on-device. The app is free if anyone wants to try it. Happy to answer questions…

A researcher trained a 125-million parameter transformer to autocomplete piano performances on-device, achieving real-time results (~108 notes per second on an iPhone 15). The key to success lay in the right MIDI representation, aggressive data cleaning, and applying DPO post-training. The project began over a year ago with the goal of using AI to autocomplete piano music, similar to GitHub Copilot but for piano.

MIDI files represent music as a sequence of events, including key presses, releases, sustain pedal changes, and instrument switches. The author primarily focused on piano-like material, removing or reducing other events. They transformed MIDI events into a discrete sequence suitable for the transformer model, choosing a representation that represented time shifts using delta_onset, chords as multiple notes with zero delta_onset, and avoiding note-off drift by ensuring note duration was explicit.

The model was built around five categorical fields for each note, including pitch, delta, duration, sustain, and velocity, with separate output heads for each field. The transformer backbone processed the notes once, rather than once per field, improving efficiency. The model also incorporated sustain duration preprocessing, extending note duration when the sustain pedal was down.

The author searched for MIDI datasets, mostly focusing on older classical music in the public domain. Cleaning and selecting the data proved more important than simply increasing dataset size. The training objective was cross-entropy over the five output heads, enabling the tracking of accuracy for pitch, duration, and velocity separately. However, the cross-entropy loss had limitations, as music continuation has no single correct answer. Augmenting the data was crucial due to the imperfect nature of live input MIDI files.

Written by urgent.news from Hacker News Best's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at simedw.com →

More in AI

Connecting Strava to Claude via MCP

Como conectar o Claude ao Strava (passo a passo para iniciantes) O Strava lançou um conector oficial baseado em MCP (Model Context Protocol), que permite ao Claude ler os dados da sua conta Strava…

More from Thursday 20 August →