Navigation

    Voting Theory Forum

    • Register
    • Login
    • Search
    • Recent
    • Categories
    • Tags
    • Popular
    • Users
    • Groups

    Why, unfortunately, the Condorcet criterion may not actually matter that much

    Voting Theoretic Criteria
    2
    5
    50
    Loading More Posts
    • Oldest to Newest
    • Newest to Oldest
    • Most Votes
    Reply
    • Reply as topic
    Log in to reply
    This topic has been deleted. Only users with topic management privileges can see it.
    • C
      cfrank last edited by cfrank

      Over the past few weeks, I’ve been thinking a lot about sincere Smith compliance under what seems to be its main strategic adversary: burial.

      I’m increasingly coming to the conclusion that formal Smith compliance may be much less informative about strategic behavior than I previously assumed. A Smith-compliant method will, of course, continue to elect from the Smith set of the ballots actually cast. But strategic burial can change that reported Smith set. So a method can remain perfectly Smith-compliant while nevertheless failing to elect a candidate who would have been the Condorcet winner (or more generally Smith compliant) under sincere preferences.

      Put differently, there seems to be an important distinction between formal Smith compliance and strategic preservation of the sincere Smith set.

      I’ve been exploring this through a large number of LLM-assisted simulations of adaptive strategic behavior. In the simulations, strategic blocs are allowed to alter their ballots freely and adapt in response to the strategies of other blocs. These are obviously exploratory tests and should be taken with a grain of salt, as I may be out of my element.

      In any case, conditioning specifically on elections in which a unique sincere Condorcet winner exists, I’ve found that even fairly sophisticated two-round mechanisms designed specifically to counter burial do not seem to improve the probability that the sincere CW actually wins by very much. The best variants I’ve tested have been around 72% under adaptive strategy.

      What surprised me is that this is very close to the performance of much simpler one-shot methods under the same strategic conditions. IRV, for example, elected the sincere Condorcet winner about 70% of the time. Stable Voting and Ranked Pairs were roughly in the high-60% range, and differences between the better-performing methods were often pretty small.

      So I don’t mean that IRV is somehow “more Condorcet compliant” than Ranked Pairs or Stable Voting. Obviously it isn’t in the formal criterion sense. Rather, it seems that formal Condorcet/Smith compliance did not translate into dramatically greater preservation of the sincere Condorcet winner once strategic behavior was introduced.

      This seems important, because we often talk about Smith compliance as though it provides strong protection against electing the “wrong” candidate. But the guarantee of Smith compliance applies only to the Smith set induced by the submitted ballots. If strategic voting substantially changes the pairwise structure, that guarantee can diverge from the thing we may actually care about: whether a candidate who would beat everyone else under sincere preferences wins the election.

      That has made me wonder whether formal Smith compliance is less useful as a practical discriminator between methods than I had thought. Maybe a more relevant question is: how difficult, risky, or dependent on coordination is it for strategic voting to displace the sincere Condorcet winner? And that’s a different property from formal Smith compliance itself.

      In fact, a formally non-Smith-compliant method might preserve the sincere CW about as often as a formally Smith-compliant method under strategic conditions, since a Smith-compliant method might remain perfectly compliant with its reported ballots even after burial has manipulated the pairwise structure away from sincerity.

      I’m still testing this, and I’m not an expert in this area, so I wouldn’t put too much weight on the values yet. But the general pattern has been pretty consistent: once sufficiently adaptive strategic behavior is allowed, a number of very different methods seem to converge toward fairly similar rates of sincere-CW preservation—namely, in this case, ~65-75%.

      What do you think of this? Personally, it makes me more interested in IRV, since it consistently outperformed many other methods under strategic pressure. I haven’t looked into the sincere bottom-Smith situation yet, but that’s another thing to consider.

      At the same time, it makes me feel that prioritizing some criteria that are incompatible with formal Smith may be a wise decision, such as participation.

      Shout out to this post: https://www.votingtheory.org/forum/topic/620/score-voting-is-king-condorcet-not-so-much/2

      cardinal-condorcet [10] ranked-condorcet [9] approval [8] score [7] ranked-bucklin [6] star [5] ranked-irv [4] ranked-borda [3] for-against [2] distribute [1] choose-one [0]

      1 Reply Last reply Reply Quote 0
      • masiarek
        masiarek last edited by

        AI Claude - it helped me to create this response:

        The distinction you're drawing is real, and I think it's underrated: formal Smith compliance is a property of the ballot→winner map, so it can't say anything about sincere preferences the method never sees. Worth separating out explicitly.

        Two things I'd push on before the IRV conclusion, though.

        First, the baseline. Ranked Pairs and Stable Voting elect the sincere CW 100% of the time on sincere ballots; IRV doesn't. So a shared ~70% isn't a tie — for the Condorcet methods the entire shortfall requires someone to have lied, while for IRV a chunk of it is the method missing with everyone honest. Which raises the diagnostic question: what's your profile generator? If IRV is reaching 70% under adaptive strategy, the model is probably impartial culture. That's where IRV always looks best (in my own runs it beats STAR there at 3 candidates, 96.7% vs 89.7% — happy to share) and it's also the model that manufactures cycles at rates no real electorate shows. Under 1-D spatial at 7 candidates, IRV elects the sincere CW 47% of the time with fully sincere ballots. Model choice swings IRV ~50 points in my measurements — wider than the entire 65–75% band you're describing as convergence.

        Second, the strategy space. Burial is the one attack IRV is immune to by construction (later-no-harm — it never reads your lower ranks while your top choice is alive), so testing burial and finding IRV robust is close to definitional. IRV's real exposure is compromising and center squeeze. Wolk/Quinn/Ogren's PVSI measures per-strategy incentive and finds IRV's favorite-betrayal incentive positive (~3%) while Smith/Minimax is disincentivized across every strategy tested — the opposite of convergence. Running adaptive compromising symmetrically would be the strong version of your test.

        One methodological question that I think drives your whole result: what are the strategic blocs maximizing? If the objective is "displace the CW," an adaptive search will always find burials. If it's the bloc's own expected utility, they'll often decline, because burial into a margin-based Condorcet rule frequently backfires and elects the buried candidate. Without backfire priced into the decision rule, 72% is an upper bound on damage rather than a prediction of behavior.

        Related: "Smith-compliant" may be too coarse to be your independent variable. On the Alaska 2022 numbers, the same burial attack succeeds or fails depending purely on the completion rule — margin-based rules shrug it off, a Hare/runoff completion falls for it. Finding that everything in the bucket lands at ~70% might be telling you the bucket is wrong.

        Also worth noting your own numbers don't quite support the conclusion: best variants 72%, IRV 70%, Ranked Pairs/Stable Voting high-60s. That's a wash within noise, not IRV outperforming — and it's a wash before correcting for IRV's sincere-ballot penalty.

        None of which is a defense of Smith compliance as a strategy guarantee — Gibbard–Satterthwaite already forbids strategy-proofness for everyone, so that was never on the table. But I'd land somewhere different from you: the useful question isn't the hit rate, it's the price — how big a coordinated bloc, how good the polling, how bad the backfire, and does it leave fingerprints. On that measure burial is a heist and center squeeze is a Tuesday.

        Would genuinely like to see the code and the generator. If you publish it I'll run it against my harness.

        C 1 Reply Last reply Reply Quote 1
        • C
          cfrank @masiarek last edited by cfrank

          @masiarek thank you, I appreciate your feedback.

          First, to clarify the simulation: ChatGPT used lightweight PSRO-like strategic agents representing coordinated factions. Each faction’s objective was to maximize its members’ mean underlying utility for the elected candidate, not to defeat the sincere Condorcet winner. An evolutionary best-response procedure searched over strategically available ballot policies against the current mixture of opposing strategies. This was done in spare time with ChatGPT, so I don’t have the code base yet, but can share once I dig it out.

          The three largest sincere-favorite factions were strategic agents. Their policies controlled strategic participation/sincerity and candidate-specific rank and score offsets. I used 9 candidates and 71 voters per election.

          Also, the benchmark was not impartial culture. I was using synthetic stress profiles including center-squeeze, polarized, clone-heavy, and ring-type electorates. So your broader point about generator dependence still applies, but the ~70% IRV result wasn’t coming from IC.

          So, in that respect, a manipulation that displaced the sincere Condorcet winner but produced a worse outcome for the manipulating faction would count against that strategy. That said, model dependence is definitely a serious limitation.

          Second, I did test a fairly broad range of completion rules for Smith-compliant methods rather than treating “Smith” as a single method. These included Ranked Pairs, B2R, IRV, Approval, Score, Benham, and several other standard or experimental completions. I also tested variants using fresh second-round ballots, including cases where the second-round completion was itself Smith-restricted. I may certainly have missed some more obscure possibilities, but the convergence did not seem specific to one particular Smith completion.

          The main failure mode was often burial or burial-like pairwise distortion: the sincere Condorcet winner had already been pushed out of the Smith set before the completion rule was applied, so at that point changing the completion rule could not rescue it, although different completion rules might change some strategic incentives and definitely impact results once the sincere CW makes it into the Smith set.

          I agree with your assessment in an important respect. These simulations assume unusually capable, coordinated strategic factions. They are basically searches for strategies that sophisticated actors could discover, rather than models or predictions of what ordinary voters would actually do, discover, or successfully coordinate.

          There is also a normative question here. A system can perform relatively well once voters behave ruthlessly and strategically while performing worse under sincerity, but it is not obvious that this is desirable, and it is unlikely to be descriptively realistic. Obviously, effective strategy has informational, computational, coordination, and cognitive costs that many voters will not pay—or may not be able to pay equally.

          I did some follow-up experiments after reading your response where I restricted behavior to be more “human-like” in the loose sense that strategic agents used simpler heuristic strategies, and I varied the fraction of strategic voters. The results were quite different: conditioning on elections with a unique sincere Condorcet winner, Ranked Pairs elected that winner about 97%, 90%, 81%, and 74% of the time when 25%, 50%, 75%, and 90% of voters, respectively, were strategic. IRV remained around 63–65% across those conditions.

          You can see these plotted here.

          So I think the distinction you’re drawing is important, and I largely agree. The earlier convergence between IRV and Condorcet methods seems to be a result about sufficiently powerful adaptive strategic behavior, not about strategic voting in general, or in human circumstances. Under simpler and more heterogeneous strategic behavior, the sincere advantage of Condorcet methods can remain very large.

          However, I don’t think the highly adaptive case is therefore unimportant. As AI tools become increasingly capable and accessible, the informational and computational costs of identifying sophisticated election strategies may decline substantially. So even if the adaptive simulations are poor descriptive models of present-day individual voter behavior, they may still be useful as stress tests of how a voting method behaves when strategic optimization becomes easier.

          cardinal-condorcet [10] ranked-condorcet [9] approval [8] score [7] ranked-bucklin [6] star [5] ranked-irv [4] ranked-borda [3] for-against [2] distribute [1] choose-one [0]

          1 Reply Last reply Reply Quote 0
          • masiarek
            masiarek last edited by

            Again AI Claude assisted 🙂

            @cfrank This is a really generous reply, and it corrects me on two things I got wrong, so let me start there.

            Your objective was already utility-maximizing for each faction, not CW-displacement — so the question I made the most noise about was one you'd already answered. And your generator wasn't impartial culture, it was stress profiles. Both of my sharpest points were aimed at problems you didn't have. Apologies for assuming.

            I spent some time with your plots, and I think they make your own case more precisely than the summary does — including in one place where they cut against your original conclusion.

            The crossover only happens at the PSRO endpoint. Reading off your panels (so ±1 point): at 0% strategic the Smith methods are at 100% CW rate and 0.992 normalized utility against IRV's ~59% and 0.944. At 25/50/75/90% they're at ~90/81/74/70% and 0.987→0.955, against IRV flat at ~61–64% and ~0.937–0.944. So on both the hit rate and the utility measure, Smith methods dominate IRV at every level of strategic participation you tested short of the endpoint — including 90%, where nine voters in ten are strategic. That's a much stronger statement than "the sincere advantage can remain large," and it's your data, not mine.

            Your "outside sincere Smith rate" panel is the cleanest thing in the set. IRV sits flat at roughly 36–40% at every level, including 0%. It elects outside the sincere Smith set about 40% of the time with nobody lying at all. That flat line is the whole reason the convergence appears: strategy has very little left to take from IRV, because IRV has already spent it. The Smith methods start at 0% and are driven up to 25–30% by strategy; IRV starts at 40% and stays. Those are different failures wearing the same number.

            Two smaller things. First, I think there may be an off-by-one between your prose and your plot: you quote Ranked Pairs at 97/90/81/74 for 25/50/75/90%, but the plot reads roughly 90/81.5/73.5/70 at those positions — your sequence looks shifted one notch. Doesn't change the shape, but at 90% strategic it's ~70%, not 74%.

            Second, and I think this matters for the conclusion: at the PSRO endpoint your best method is IRV + fresh runoff, not IRV. It leads on all three panels — ~69% CW rate, 0.964 utility, and the lowest outside-sincere-Smith rate of any method at that endpoint (~31%, while plain IRV is at ~39.5% and Ranked Pairs is at ~56%). So even taking the adaptive regime entirely at face value, the finding isn't "IRV is robust to sophisticated strategy." It's "a second round on fresh ballots is robust to sophisticated strategy" — a claim about two-round structure rather than about Smith compliance, and one that would apply just as well to score-plus-runoff designs. I'd be curious whether that holds up as you vary the runoff's ballot.

            On your mechanism — the CW being pushed out of the Smith set before completion runs. That's a sharp claim and I hadn't separated it out, so I measured it. Single coordinated bloc, burial, rational (it only submits if it beats voting honestly), every challenger tried, swept over 3/5/7/9 candidates at your 71 voters. Ejection is the minority regime everywhere: 19–36% of successful burials remove the CW from the reported Smith set, and the other two-thirds to four-fifths leave the CW sitting inside it, where the completion rule still decides. Interestingly the ejected share barely moves with field size (23%→26% in 1-D from 3 to 9 candidates) even as raw displacement climbs from 20% to 88%. A wider field makes burial much easier without much changing where it lands.

            The obvious limit: that's one bloc doing a plain bury-to-last, not several factions best-responding over rank and score offsets. Ejection is clearly something a stronger search buys. Which I think locates our remaining disagreement precisely and answerably: how much of the ejection rate is purchased by adaptive multi-faction optimization, over and above single-bloc burial? If PSRO is ejecting at 60–70% where a naive bloc ejects at 25%, that's a real and quotable finding about what adaptivity does, and it would be worth a paper on its own. Code's here if it's useful: https://masiarek.github.io/star-voting-library/07_Concepts/topics/compliance_vs_strategic_preservation.html

            On the AI point — I think this is the most interesting thing in your post and I don't want to wave it away, because the direction is right. Three things give me pause about how far it goes:

            Computation isn't the binding constraint. In my runs a successful burial needed 34–40% of the electorate to rank someone they genuinely like dead last. AI can tell you that's the optimal play; it can't get 40% of voters to cast it on trust. That cost is social, and it's the one that doesn't fall with compute.

            The information required is about opponents, not preferences. PSRO converges because it iterates against a live opponent mixture, observing what the other factions actually do. A real electorate votes once, without seeing anyone's final policy. Better tools don't close that gap — they arguably widen the variance, since everyone is now optimizing against a guess about everyone else, and a burial aimed at the wrong equilibrium is exactly the one that backfires.

            Cheaper attacks make cheaper detection, and that cuts toward pairwise methods. A successful burial's signature is a cycle appearing in a race whose pre-election pairwise polling showed a clean head-to-head winner. The thing detection needs is a published pairwise matrix — which is precisely what Condorcet methods emit as a byproduct and what IRV doesn't. If the threat model is "strategic optimization gets cheap," I'd want the method that publishes the most auditable structure, not the least.

            None of which touches your core point, which I think is right and which I've now written up: formal compliance is a property of the cast ballots, "the sincere winner wins" is a property of the electorate plus its behaviour, and treating the first as a guarantee of the second is sloppy. That's a correction advocates on my side of this should absorb rather than argue with. I just don't think it demotes Smith compliance — your own 0–90% panels are about as strong an argument for it as I've seen.

            C 1 Reply Last reply Reply Quote 1
            • C
              cfrank @masiarek last edited by

              @masiarek you're right, the "97/90/81/74" sequence was from a different experiment. The plots I shared have the accurate rates for the relevant experiments.

              Your point about information is also important, because PSRO involves repeated optimization against mixed opponent strategies with effectively complete information about the modeled strategic environment, and a real election is definitely not like that. Present elections are probably better represented by the bounded-strategy region than by the PSRO endpoint, which does strengthen the practical case for Smith compliance. But declining computational and informational costs could move real strategic behavior toward that endpoint.

              Also to clarify, the initial post was based entirely on the PSRO adaptive-strategy results. The later "human-like" bounded-strategy experiments were prompted by curiosity based on your feedback, and they changed the interpretation in the opposite direction. Unless we assume extremely capable strategic agents with very strong coordination and information, the PSRO results probably aren't descriptively realistic except as an approximation to a highly sophisticated strategic limit. It is interesting and somewhat unfortunate, though, that most methods seem basically unable to withstand the kind of strategic pressure PSRO simulates, at least in terms of preserving the sincere Smith structure.

              Your last point about the fresh runoff aspect is also correct. I did also try several Smith-compliant methods with fresh runoffs during exploration, but those should be compared again systematically. In my exploratory simulations, most fresh runoffs showed only marginal improvements under PSRO, except for IRV vs runner up. There may be some interesting theoretical reasons for this, see Durand et al. 2026, "Super Condorcet Winners and Limit Coalitional Manipulability of IRV"--this uses Impartial Culture. I did not try fresh-runoff Smith methods in the later "human-like" analysis, so that would definitely be interesting to test as well.

              cardinal-condorcet [10] ranked-condorcet [9] approval [8] score [7] ranked-bucklin [6] star [5] ranked-irv [4] ranked-borda [3] for-against [2] distribute [1] choose-one [0]

              1 Reply Last reply Reply Quote 0
              • First post
                Last post