Problem statistics: how each metric is computed — MOJ docs

Problem statistics: how each metric is computed

Translation note. This manual is a translation of the Portuguese original. The command-line tools (moj, moj-contest, moj-comp) print their messages in Portuguese, and the command examples below are identical to the original.

This document explains, section by section, how the Problem statistics page of Free Training (/treino/problema/stats/?id=<problema>) calculates what it shows. It includes the statistical decisions and the honest limitations of each number. The source of truth is the endpoint GET /treino/problem-stats (full contract in API.md). This page describes the semantics.

Where the data comes from

Each training account keeps its own submission history (1 line per submission: problem, language, verdict, and time in epoch). The statistics of a problem are the aggregation of all the lines of all the users for that problem. They come only from Free Training: submissions made in class contests are not included.

Summary

Metric Calculation
submissions total number of history lines of the problem
attempted distinct users with ≥1 submission
solved distinct users with ≥1 accepted submission
solve it (per user) solved ÷ attempted — the per-user rate
per-submission rate accepted submissions ÷ total submissions — measures how much people fail while trying; it does not define the difficulty
subs / user submissions ÷ attempted
difficulty label from the per-user rate: ≥90% very easy · ≥70% easy · ≥50% medium · <50% hard · no users who attempted = new
dirt (submissions of the solvers up to the 1st AC − ACs) ÷ (those submissions). It is the ICPC resolver metric, the same one as in the contest statistics. High = the problem punishes mistakes.

The difficulty has a single source in the system (lib/difficulty.sh on the server, shared/difficulty.js on the web). The search, the suggestion, the profile, the contest draw, and this page read the same key. Before, this page labeled by the per-submission rate. As a result, the same problem showed "easy" in the search and "hard" here (issue #30). The per-submission rate stays on the page as a number, with the correct name.

Difficulty percentile against the archive

The "X% of the archive is easier than this" card compares the per-user success rate (solved ÷ attempted) of this problem with the rate of all the public training problems (the same base as the problem list).

Facts

Timeline

Activity calendar

How they solve

Running time (accepted submissions)


Endpoint contract (fields and formats): API.md, route /treino/problem-stats. The display of verdicts follows the central policy of the platform (single source lib/verdict.sh — verdicts are never translated).