PulseAugur
EN
LIVE 10:36:27

User trained small LLM with custom dataset 1-2 years ago

A user reflects on training a small LLM approximately one to two years ago using a custom dataset comprising their Mastodon posts, personal writings, code, and Q&A factoids, supplemented by open datasets like SQuAD. The user notes that the initial model was not highly intelligent but improved over time, with training primarily conducted on personal hardware and partially on a cloud platform due to cost considerations. AI

RANK_REASON The item describes a personal anecdote about training a small LLM, which does not constitute a significant industry event, research breakthrough, or product release.

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

User trained small LLM with custom dataset 1-2 years ago

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Meme
The item describes a personal anecdote about training a small LLM, which does not constitute a significant industry event, research breakthrough, or product release.
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    thinking about that time about 1-2 years ago when I trained a small LLM with my own non-scraped dataset of my own data (Mastodon posts, some Q&A factoids, some

    thinking about that time about 1-2 years ago when I trained a small LLM with my own non-scraped dataset of my own data (Mastodon posts, some Q&A factoids, some of my own code, and things I wrote) + some existing open datasets like https:// rajpurkar.github.io/SQuAD-expl orer/ ..i…