What actually lowers risk, ranked by how strong the effect is and how good the evidence for it is.
What has to be decided first
How the two axes — effect size and evidence strength — are combined without letting a large claimed effect on weak evidence float to the top.
Tracked as ticket 008-ranking-model