Vinyl Taste Profile & Discoveries¶
An analysis of my record collection — what I listen to, where it clusters, and what the algorithm has surfaced that I don't own yet.
User: barneypinkerton Collection: 194 releases | Wantlist: 135 releases
1. Collection DNA¶
What styles, labels, and eras define my collection. Wantlist items are weighted 1.5× – since I haven't bought many records recently but have always maintained my wantlist, the wantlist is a better representation of my current taste preferences than existing purchases in my collection.
2. The Mean Record¶
I've used my local MP3 collection of dance music – gathered over the years via Bandcamp, Beatport, and Soulseek – to build a digital portrait of my taste. Every track has an audio embedding – a 1280-dimensional fingerprint computed by Essentia's EffNet model trained on the Discogs catalogue. Each song is then placed in a shared multidimensional space based on its characteristics. The centroid is the average of all those fingerprints. The track below sits closest to it: the most typical sounding song in the collection.
For me this holds true, 'Tripped' by Wheelman is a personal favourite of mine and has everything I love in a dance track: lush synths, skippy drums and a trippy vocal sample, a good sign the model has read my taste accurately.
Collection map¶
Below, every track is placed by its true distance from the centroid, measured directly in the full 1280-dimensional embedding space (no lossy 2D projection). The centroid — the true mathematical average, not an actual track — sits at the very centre; every real track spirals outward the less it resembles it. The angle around the spiral carries no meaning — it's just spacing so points don't overlap. Only the distance from centre, and colour, reflect real similarity.
3. Discoveries¶
Below are some of my favourite records that have been surfaced by the recommendation pipeline so far.
Candidates come from two separate routes. The first, affinity, works directly off the connections in my collection: if I own records on label X and wantlist records on label Y, it looks for what sits at the intersection of those scenes (maybe there's an artist that's released on both those labels that I don't own in my collection) things I'm likely to know, or should. The second, discovery, ignores those direct links entirely and instead casts a wide net using the overall DNA of my collection – its style and era profile – to surface records with no direct tie to anything I own.
Both routes are then passed through the same filter: the EffNet audio model, which checks whether a candidate actually sounds like the rest of my collection and re-ranks accordingly. The pipeline will analyse the YouTube audio clips embedded into the Discogs release page and cross-reference with the embeddings of my local MP3 collection.
So affinity picks are ones the graph already half-expected and the audio confirms; discovery picks are ones the audio vouches for even though nothing in my library links to them directly. These have been the more interesting recommendations.
Favourites so far¶
Bouncy synths and heavy bass





