Use PR AUC when the positive class is rare and the business goal is to find that class reliably. PR AUC focuses on precision and recall, so it is better suited to problems like disease detection or fraud screening, where missing minority events matters more than rewarding true negatives. If class balance and negative-class performance matter equally, ROC AUC is usually the better fit.
Why PR AUC Fits Heavily Imbalanced Classification
pr auc is the better primary metric when the positive class is rare and the cost of missing it is high. It measures the trade-off between precision and recall directly, so it tells you whether the model is finding the minority class without flooding operators with false positives. That makes it more informative than accuracy, and often more useful than ROC AUC, in skewed detection problems.
When the negative class dominates, ROC AUC can look strong even if the model performs poorly on the rare class. PR AUC removes most of that illusion because true negatives are not doing the work of making the score appear good. In practice, that means PR AUC is a stronger fit for screening, alerting, and triage workflows where the minority event is the object of interest.
PR AUC is also a better choice when the business decision is about ranking a limited review queue. If teams only have capacity to investigate a small fraction of cases, they need to know how many of the flagged items are actually useful and how many true positives the model recovers. That is exactly the operational question precision and recall answer together.
For a practical comparison, the right metric depends on what the downstream decision values most. If false negatives are especially costly and the positive class is the target of the system, PR AUC should usually be the headline metric. If the problem requires balanced treatment of both classes, or if true negatives are part of the core success criterion, ROC AUC is often the more stable summary.
What PR AUC Does Not Tell You on Its Own
PR AUC is a summary metric, not a complete decision rule. A model with a strong PR AUC can still be poor at a specific operating threshold if the threshold is set badly for the workload or the available review capacity. Teams still need threshold analysis, calibration checks, and confusion-matrix review at the point where the model will actually be used.
Class prevalence also matters. PR AUC is sensitive to the base rate of the positive class, which is part of its value in imbalanced settings, but it also means scores are not always easy to compare across datasets with very different prevalence. If the target population changes, the same PR AUC can reflect a different level of practical usefulness.
That is why metric choice should follow the decision context, not just the training data shape. If the model is intended to prioritise rare events, PR AUC is the right primary lens. If the question is overall ranking quality across both classes, ROC AUC may complement it. For heavily imbalanced classification, the best practice is usually to report both, but to optimise against PR AUC when minority-class retrieval is the business objective.
How ML Teams Should Choose the Metric
The cleanest decision rule is simple: use PR AUC when the positive class is rare and the cost of missing positives is materially higher than the cost of additional false alarms. Use ROC AUC when class balance matters more, or when you need a metric that is less tied to prevalence and more reflective of general ranking separation.
Teams should also decide whether the metric will be used for model selection, model monitoring, or stakeholder reporting. A model that wins on PR AUC in evaluation may still need a different operational threshold in production, while a model that looks acceptable on ROC AUC may fail the actual review workflow because precision is too low. The metric should match the decision, not just the dataset.
For heavily imbalanced problems, the most useful habit is to anchor evaluation to the minority-class workflow: how many true positives are recovered, how much manual review is created, and whether the model meaningfully improves triage. That framing prevents teams from overvaluing metrics that are mathematically elegant but operationally misleading.
Practitioner Guidance
What to prioritise: Treat PR AUC as the primary selection metric when the rare class is the target outcome and review capacity is limited. Pair it with a threshold-specific precision and recall view so you can see whether the model is useful at the point of deployment, not just in aggregate.
What to verify: Confirm that the positive class rate in evaluation matches the real operating population closely enough for PR AUC to be meaningful. If prevalence shifts materially, re-check the model against the new base rate before trusting the score.
Common mistake: Do not use ROC AUC alone to justify a model for rare-event detection. A model can separate classes well in ROC space and still perform poorly where it matters, especially if precision at the usable threshold is too low.
Practitioner takeaway: For heavy imbalance, choose the metric around the decision you are trying to support, and if the goal is to reliably surface the rare positive class, PR AUC should usually drive model selection.
Related resources from NHI Mgmt Group
- How should teams choose F1 score instead of accuracy for imbalanced classification problems?
- How should ML teams choose a classification threshold when false positives are expensive?
- How should security teams choose between data classification tools for cloud and AI estates?
- What do security and data teams get wrong about imbalanced classification?