Table of Contents
Inside the Standard Farm & Stream finder
Why does one map enter the Farm pool while another enters Stream? The osu!standard finder studies two different things: PP efficiency in recorded player scores, and rapid tapping patterns in the map itself. A map can qualify for both.
This article follows the code from collection to ranking to the bot's next-map selection. For playing rather than implementation details, start with .farm, .stream, .variety and .pool.
| Scope | osu!standard; Mania uses separate research and style detection |
|---|---|
| Research reference | Standard finder source bundle dated 2026-09-27: Farm Score v2.2, stream_v3, stream_q4 |
| Live integration reference | PoolRotater v6.6.17, reviewed 2026-09-29 |
| Status | Experimental research; the live bot reads an existing catalog |
The reviewed finder bundle is staged on the VPS. Its presence does not mean a new collector is running, that its latest report has replaced the live catalog, or that every research qualifier is enforced in lobbies. The research pipeline and live integration are distinguished below.
The two paths
Player top-100 scores + map metadata Cached .osu hit objects
| |
Formula + realized PP models Rapid, even circle runs
| |
Farm Score v2.2 Stream Score v3
| |
farm_report stream_features
+--------------+---------------+
|
Ranked maps + room filters + source mix
|
Freshness + player feedback + pick
Farm asks: does this map return unusually much PP compared with similar recorded performances and players?
Stream asks: does this map contain enough rapid, evenly timed circle sequences to have a stream identity?
Farm Score is a relative ranking. Stream Score is a structural score between 0 and 1. The bot keeps their rankings separate; it does not average the two numbers.
1. Collecting the evidence
collect.py samples players across PP brackets, rather than studying only the top of the leaderboard. Its default bracket quotas total 23,250 target players. That is a collection target, not a claim about the completed database.
For each successfully refreshed player, the collector stores their top 100 scores and fetches the corresponding map metadata. A successful refresh replaces that player's previous score snapshot, so a score that has fallen out of the top 100 is removed. Checkpoints support resuming collection.
Country diversity caps are relaxed when brackets are hard to fill. The reviewed source also excludes country code RU from collection and model score queries. These choices affect coverage: this is a filtered research sample, not a census of osu! players.
What the sample cannot tell us: top-100 scores are selected successes. They omit most attempts, failures, practice sessions and lower-value plays. A map absent from the data may simply be underrepresented. A low model score does not mean a map is bad.
The report can optionally filter scores using the source's REWORK_DATE boundary of 2026-07-26. farm_report() defaults to all eligible recorded scores, not that date-filtered subset. If a requested subset has too few scores, the model can warn and fall back to all scores. Each fit requires at least 2,000 usable observations.
2. Farm: two views of PP efficiency
features.py fits two histogram gradient-boosting regression models to logarithmic PP values from recorded scores. These are learned approximations, not an implementation of the official PP formula or a prediction of your next profile gain.
| Model | Inputs | Question |
|---|---|---|
| Formula efficiency | Star rating, accuracy, misses, combo ratio, mod flags | Given this headline difficulty and performance, is the recorded PP unusually high? |
| Realized efficiency | Player's total PP, star rating, playable length, mod flags; player PP and length are log-transformed | What do similarly skilled players actually extract from maps of this difficulty and length? |
Both models include HD, HR, DT, FL, EZ, HT and NF flags. NC is normalized into the DT modeling bucket. The stored map star rating is the API's no-mod rating; a binary mod flag cannot capture every change in difficulty.
The first model deliberately omits features such as AR, OD, CS, length and object counts. This leaves differences for the residual to reveal, instead of teaching the model to reproduce every part of the PP calculation. The second deliberately omits accuracy, misses and combo: practical performance belongs in the outcome it is trying to measure.
Residuals: the signal left over
For each score:
residual = log(recorded PP) - model prediction of log(PP)
A positive residual means more PP than that model expected; a negative residual means less. The code averages residuals per beatmap and, by default, requires at least five observations.
It then uses empirical-Bayes shrinkage: an uncertain map average is pulled toward zero. More consistent evidence allows more of the original estimate to survive. Afterwards it centers map estimates within their dominant mod bucket, reducing systematic offsets between buckets such as NM and DT. This is a per-map dominant bucket, not a separate final ranking for every mod combination.
_robust_z() standardizes the two resulting signals using their median and median absolute deviation, with fallback scales for degenerate data and a clip at ±6. This makes them comparable without letting a few extreme maps set the scale.
Combining the signals
The default v2.2 calculation is:
base = 0.65 * formula_z + 0.35 * realized_z before_calibration = base - disagreement_penalty FarmScore = before_calibration - SR_local_adjustment
The disagreement penalty applies only when the shrunken formula residual is positive and the shrunken realized residual is negative. A map can appear to overpay while similarly skilled players still struggle to extract that return. The default penalty coefficient is 0.30; it multiplies the positive and negative strengths, each divided by its robust scale and clipped to 0–6. These strengths preserve zero as the sign boundary, unlike median-centered z-scores.
The star-rating calibration subtracts a smooth estimate of the typical score around the map's SR. Default settings use 0.25★ bins, a 1.25★ Gaussian bandwidth and a 300-map support prior. Bin medians use square-root count weighting; effective support controls how much of the estimated local baseline is subtracted. At 300 effective supporting maps the adjustment is half-strength. This removes a local offset; it does not rescale each difficulty band.
For illustration only: formula_z = 1.4 and realized_z = 0.6 give a base of 1.12. With no disagreement penalty and an SR adjustment of 0.15, FarmScore is 0.97. That is neither 97% confidence nor 0.97 PP.
Confidence and diagnostics
Confidence is reported separately:
confidence = observations / (observations + 20)
Five observations give 0.20; 20 give 0.50; 80 give 0.80. Labels are explore below 0.50, supported from 0.50, and high_conf from 0.80. This confidence is not a bonus or multiplier in FarmScore; low-sample estimates are already affected by shrinkage.
The report also contains popularity/adoption, profile impact, forgiveness and performance diagnostics. Popularity, age, raw score count, farmer clustering and forgiveness do not contribute positive ranking terms in v2.2. Counts still matter for estimation and confidence, and pass count can be an eligibility filter. A popular map may rank highly because of its evidence, but popularity itself is not the farm reward.
Requested SR, length, pass-count and mod filters are applied before selecting the report's top rows. The report saves farm_report in SQLite and can export CSV results.
3. Stream: reading the map
rotation.py implements find_tap_runs() and analyze_stream_file(). stream_farm.py reuses this classifier and can store its results alongside the Farm database.
The parser reads hit-object timestamps and positions from a cached .osu file. A tap run consists of consecutive circles with rapid, nearly even intervals. Sliders and spinners break runs. The default interval window is 45–135 ms, with 15% tolerance against a rolling median of recent intervals.
A run's stream BPM is 15000 / interval_ms, the quarter-note tapping equivalent. This is derived from hit-object timing, not simply copied from the song's metadata BPM. The classifier also measures the distance between circles.
| Run length | Classification |
|---|---|
| 3–4 notes | Small bursts; recorded separately |
| 5–7 notes | Short streams |
| 8+ notes | Sustained streams |
| 12+ notes | Long streams; also counted among sustained streams |
Coverage ratios divide notes by the map's playable-object count, including sliders and excluding spinners. The classes therefore describe circle patterns within the whole playable map, not only within its circle count. Long-stream and sustained ratios overlap; adding them would double-count notes.
Two ways to establish stream identity
stream_v3 takes the larger of two scores. Every component below is capped to the 0–1 range before weighting.
| Sustained path | Weight | Component reaches full strength at |
|---|---|---|
| 8+ note coverage | 0.35 | 30% of playable objects |
| 12+ note coverage | 0.25 | 18% of playable objects |
| Longest sustained run | 0.20 | 32 notes; strength starts at 8 |
| Sustained run count | 0.10 | Four runs |
| Long run count | 0.10 | Three runs |
| Repeated-short path | Weight | Component reaches full strength at |
|---|---|---|
| Coverage of all 5+ note runs | 0.35 | 12% of playable objects |
| Number of 5+ note runs | 0.30 | Four runs |
| Total notes in 5+ note runs | 0.20 | 24 notes |
| Longest 5+ note run | 0.15 | Eight notes; strength starts at four |
StreamScore = max(sustained_score, repeated_short_score)
For illustration, four separate five-note runs in a 400-object map give 20 notes and 5% coverage. The repeated-short formula gives 0.65 even with no eight-note sustained run. That raw score alone does not prove the map meets the separate quality, spacing or lobby filters.
This allows repeated five-to-seven-note patterns to establish a map's stream identity without requiring a marathon stream. Repeated three-to-four-note bursts do not feed the same scoring path. Slow jumps will not become streams merely because the song has a high BPM.
The classifier is a timing-and-structure heuristic. It does not fully model aiming difficulty, finger control, strain, or how tiring a map will feel to a particular player. Its raw file analysis does not apply a lobby's selected speed mods.
4. The research qualifier: stream_q4
The finder also has a separate rotation catalog with a stricter Stream qualifier. Its raw StreamScore measures structure; stream_q4 additionally checks suitability.
Default hard requirements include Ranked status, the current classifier version, StreamScore ≥0.50, derived stream BPM ≥140, playable length 45–300 seconds, at least 15,000 passes, checked download availability, community rating ≥9.20, quality score ≥0.60, and mean stream spacing ≥12 pixels.
With complete metadata, community quality combines 70% normalized rating, 20% logarithmic favourites and 10% activity confidence. Rating is normalized across 8–10, favourites across 25–2,500; activity combines pass count and beatmapset play count, with map play count as a fallback. Missing rating can produce a provisional diagnostic score, but cannot pass final qualification without a real rating and availability check.
A map must also meet one of three structural paths:
- Sustained: at least 10% sustained-note coverage and 7.5% long-note coverage, plus either two long runs or a run of at least 20 notes.
- Repeated short: at least 5% coverage from all 5+ note runs, at least three such runs, and at least 18 notes across them.
- Exceptional run: a sustained run of at least 28 notes.
The exceptional-run path relaxes only the structural coverage test. BPM, spacing, length, activity, rating and availability still apply. Results store qualification, reason and qualifier version in map_quality.
These are defaults in the reviewed research catalog, not a promise that all q4 gates currently run in every live room.
5. What the live bot actually consumes
The reviewed RankedStandardMapPool in bot/pool.py reads beatmaps, farm_report and stream_features. Its Stream query does not use the research catalog's map_quality.stream_lobby_qualified flag.
| Source | Live ranking and eligibility |
|---|---|
| Farm | Existing FarmScore, descending; Ranked maps inside configured SR, maximum playable length and minimum pass count, plus ranked-age policy when configured. Dominant DT/NC report buckets are excluded. |
| Stream | Ranked maps inside configured SR, length and pass-count boundaries, with configured raw StreamScore and derived BPM thresholds. Ordered by StreamScore, longest sustained run, long-stream coverage, then pass count. |
| Mixed sources | Reserve slots for each source, merge overlapping beatmap IDs, and fill remaining capacity from available candidates. A map with both labels occupies one unique slot. |
The live Stream class defaults are 0.50 score, 140 BPM and 45-second minimum length, but room configuration can override them. Use .pool for the active settings. The Farm dominant-DT/NC exclusion concerns research buckets; it does not mean speed mods are forbidden in the room.
Ranking builds a candidate pool. It does not make the top research result play every time. Recent plays/skips, reserved queue entries, song spacing and player feedback shape subsequent selection. The reviewed final random picker uses feedback weighting within the eligible candidates, rather than directly treating FarmScore as a draw probability. How maps are selected explains the broader policy, including sparse-pool expansion.
For older compatible databases without farm_report, the code can fall back to eligible beatmaps ranked by pass count. Those candidates are not labeled Farm. Missing status evidence is not guessed into Ranked eligibility.
Limits and interpretation
- Research snapshot versus catalog: this article describes verified source code. It does not assert that the running catalog was generated with every setting shown here.
- Sample bias: player PP is a proxy for skill; top-100 scores, bracket quotas, country filters and uneven map coverage influence the learned signal.
- Validation limits: both regression fits reserve 20% of scores for validation. That is a random score split, not a guarantee of independent unseen players or maps. Reported fit quality is not proof that a recommendation will work for you.
- Changing rankings: new observations, score refreshes, model fits and calibration can change ordering. FarmScore is relative to the fitted dataset; compare it within a report rather than as an absolute cross-version promise.
- Different rulesets: osu!standard streams are circle-timing runs. Mania stream patterns and Mania Farm research use different evidence and models.
Code map and related reading
| Source | Responsibility |
|---|---|
core.py | API access, shared configuration, mod normalization and database schema |
collect.py | Player sampling, top-score refresh and map metadata collection |
features.py | PP models, per-map residual estimates and diagnostics |
analyze.py | Farm Score v2.2, disagreement penalty, SR calibration and report output |
rotation.py | .osu tap-run parser, stream_v3, quality and stream_q4 rotation-catalog qualification |
stream_farm.py | Store/reuse raw Stream features alongside the Farm database |
Live bot: bot/pool.py | Eligible source pools, separate rankings, deduplication and next-map selection |
Read selection policy, feedback, PP estimates, technical overview or the feature index. This article documents behavior; it does not trigger collection, rebuild a report or change a lobby.