SCROLL NEWS / DISCOVERY

Search headlines

Find the stories shaping the conversation.

Search mode
Refine results
Active filters Keyword All time

Results for " Inaccurate " 37 found 🔀 AI-powered Shuffle

Good science requires much more than evidence

Empiricists believe that scientific theories are true only if they make correct empirical predictions: if we can't verify these predictions, then there is no reason to expect that they tell us anything about reality. However, Professor of Philosophy at King's College David Kyle Johnson argues that this is mistaken. In response to Bas van Fraassen, who has written on this subject for the IAI, he argues that we have many reasons for believing a theory to be true, such as it being parsimonious and having large scope — qualities which have been used throughout the history of science as criteria for judging theories. In 2024, Princeton philosopher Bas van Fraassen wrote a short article for IAI News entitled “Science Does not Describe Reality: The Limits and Benefits of Explanation.” At the time, I was working on two projects (a science book and a Great Course) that, when completed, would (among other things) defend science’s ability to describe reality. I wanted to write a response to van Fraassen, but the projects demanded too much of my attention. But now that those projects are complete, I am in a position to explain the objections to van Fraassen’s thesis. Summarizing van Fraassen’s argumentIn a nutshell, his argument is this: What scientists do is make predictions based on theories and then invite others to test to see whether those predictions are borne out. “If theory T is right, result R should happen in experimental condition C.” Many think that running an experiment that shows that R happens in condition C is proof that theory T is true (and should be believed). But, van Fraassen argues, that R happened in condition C is only a reason to think that R happens in condition C. That’s the only observable, or empirical, result we can actually see, so that is all that we are really justified in believing. (The underlying theory that made that prediction still may or may not be true.) That is why he calls those who accept his position “empiricists.”Before I explain what is wrong with this position, let me explain what it gets right. What van Fraassen Gets RightFirst, a theory correctly predicting the results of an experiment is not, strictly speaking, “proof” that the theory is correct. The basic line of reasoning is this:If the theory T is true, result R should happen. Result R did happen. Thus, theory T is true.In other words:“If T then R. R. Thus, T.”This is a deductively invalid, fallacious mode of reasoning, called “affirming the consequent.” Those premises do not guarantee that the conclusion follows. “If I am in Pennsylvania, then I am in the USA. I am in the USA. Thus, I am in Pennsylvania.” Clearly that argument doesn’t guarantee its conclusion; neither does the above argument.However, science is not a deductive method of reasoning; it is an inductive one. And inductive arguments merely try to raise the probability of their conclusion, not “prove” them for certain. Thus, the fact that an experimental result doesn’t prove a theory is true is irrelevant; it’s not supposed to. That is not what scientific reasoning is about.___If the hypothesis that something exists clearly provides the best explanation for what we can observe, then the belief that it exists is what is rational.___Second, it’s also true that a single experimental result doesn’t provide enough justification to believe that a theory is true. Scientists do not come to a consensus about whether a theory is true by looking at a single experiment; they look at all the evidence to determine where the preponderance of evidence points. The mere fact that I am in the USA raises the probability that I am in Pennsylvania, but not nearly enough to conclude that I am in Pennsylvania. However, if I am also regularly encountering people wearing Eagles and Phillies hats and seeing advertisements for Founding Father tours, I probably am. In the same way, a theory successfully predicting the outcome of one experiment is not enough to conclude that it is true; but if a theory makes a wide variety of predictions that are all vindicated by multiple experiments, performed by many different groups and people, all actively trying to prove the theory false (which is what good scientific experiments are designed to do)—that is a good reason to conclude that the theory is true. In other words, if you repeatedly try to prove something false but can’t, that is a very good indication that it isn’t. The problem of underdeterminationIn reply, those whom van Fraassen calls “empiricists” often point to the “problem of underdetermination.” For any given theory that makes successful predictions about what happens in experimental conditions, there are countless other possible theories that make exactly the same predictions—and since they do, the experimental evidence in question is just as much reason to accept any of those theories as it is to accept our original one. Thus, the argument goes, we have no way of knowing whether our original theory is true, or one of the other “predictively equivalent” theories is true. SUGGESTED VIEWING How Occam's Razor changed the world With Johnjoe McFadden As an example of underdetermination, philosophers often offer this example. The evidence of fossils seems to provide reason to conclude that the universe is quite old; ancient fossils are what the “ancient universe” hypothesis predicts. But it’s also possible that God created an ancient-looking universe 200 years ago, complete with a “fake” fossil record for us to find, to make it look old. The existence of fossils is what that theory predicts too. So, the argument goes, the fossil record can’t be used as evidence that the universe is ancient; the fossil record supports the “200-year-old ancient-looking universe” theory just as much. The problem of underdetermination points out that, for any theory with supporting evidence, we could invent another completely different theory that predicts exactly the same evidence, and thus undermine that evidence’s ability to support the original theory.But this “argument from underdetermination” betrays a fundamental misunderstanding of what scientific reasoning is and how it works. To explain, let me use another rather strange (but very real and commonly discussed) philosophical thought experiment called “The Grue Problem.” The “Grue” ProblemSuppose I want to know what color emeralds are, so I observe a few, see that they all appear green, and form the theory that “all emeralds are green.” I thus predict that the next 100 emeralds I find will also appear to be green. And let’s say that happens, and so I thus conclude that my theory that all emeralds are green is true. Now, one might rightly point out that I don’t know for sure that all emeralds are green—I have not seen them all. Maybe some I haven’t found yet are blue. Fair enough. After all, we used to think that all swans were white, until we found some black ones. But let’s say that I go on a quest and collect every emerald in the universe—and, for the sake of the argument, let’s say I can do this, and can even know that I have every one of them. So, I collect them all, look at them all, and verify that they all appear green. Can’t I now rightly conclude that all emeralds are green?Inspired by the problem of underdetermination, a friend might argue that I cannot.Sure, they all appear green now—but there could be another kind of color called “grue.” Something is grue if it is green up to a certain point in time, but then spontaneously turns blue. And there could be an infinite number of “shades” of grue. Something that is one shade of grue will turn blue at noon today, but something that is another shade of grue will turn blue at noon tomorrow, etc. So yes, the emeralds appear green now; and since they haven’t turned blue yet, they are not some “past” shade of grue. But they might be some “future” shade of grue (e.g., they may turn blue at noon tomorrow). And since all future shades of grue are consistent with “being green up to this point in time,” the evidence you have is completely consistent with the notion that all emeralds are some future shade of grue, rather than green. Indeed, that they would appear green until now is what the “some future shade of grue” hypothesis predicts. So, we have just as much reason to conclude that all emeralds are some shade of grue as we do to conclude that they are green.You’re probably thinking “It’s not equally possible that the emeralds are grue; they are obviously green.” You are right! But explaining why your instinctive reaction to my friend’s argument is correct will help inform why van Fraassen’s “empiricist” is wrong and science actually does describe reality.The main mistake my friend makes is in thinking that the “empirical evidence” (of all the emeralds appearing green) is (or even should be) all I am using to draw the conclusion that all emeralds are green. It is not. I am also considering known facts about how colors typically work; things don’t usually change color, and when they do, they don’t do so suddenly, without cause, and all at once. So, even though the simple evidence of the emeralds all appearing green up to this point is consistent with them being some future shade of grue—that is even what the grue hypothesis “predicts”—it is still exceedingly unlikely that they are. The hypothesis that they are green is much more likely. Understanding science refutes the problem of underdetermination The misunderstanding about science that the “empiricist” pointing to “the problem of underdetermination” makes is similar to the one that my friend made. Science is not just about considering experimental evidence; in fact, it’s not even just about considering empirical evidence. As Ernan McMullin and many other philosophers of science have pointed out, the inference that makes science is “inference to the best explanation” (sometimes called “abduction” or “retroduction” by McMullin). While, depending on the circumstances, scientists use many different kinds of reasoning, science as a whole is, simply put, the process of considering multiple competing explanations and comparing them to certain criteria—explanatory criteria (the kinds of things explanations must be and do, by definition, to be a good explanation)—to see which explanation is the best.And these criteria do not deal only with empirical results. Yes, successful predictions raise the probability of a hypothesis, but it’s also the case that the more new assumptions a theory makes, the less likely it is to be true; every new assumption presents an opportunity for the theory to be false. (The fewer assumptions a theory makes, the simpler or “more parsimonious” it is said to be.) Likewise, the more existing evidence a theory contradicts, the less likely it is, because in order for a theory to be true, every piece of evidence it contradicts has to be somehow faulty. (The fewer established theories a new theory contradicts, the more “conservative” it is said to be.) And, obviously, the more something explains, the better explanation it is. (This is called having “scope.”) If a theory raises unanswerable questions, it’s not expanding our understanding, and thus is not wide-scoping. Scientists consider all these things when deciding which theory to accept, because the theory that adheres best to these explanatory criteria is most likely to be true.___the mere fact that a theory’s predictive success is why it “sticks around” or “becomes accepted” in the scientific community is not a reason to be skeptical about its truth.___The reason the grue hypothesis is no good is because it invents an entirely new kind of color—not just a new shade, but an entirely different category of color (a different way colors work)—and thus is not parsimonious. It raises unanswerable questions about how and why emeralds would just spontaneously change from green to blue (and thus does not have wide scope). And it contradicts what is already well established about how color works (and thus is unconservative).The theories that the problem of underdetermination points to do the same thing. Yes, the theory that God created an “ancient-looking universe” (complete with an ancient-looking fossil record) two hundred years ago is also consistent with the empirical evidence (of a fossil record) that presently exists; but that theory is monumentally unparsimonious (it requires many extra assumptions about infinite entities) and not wide-scoping (it raises unanswerable questions about why God would do such a thing). So, I can rationally conclude that it is false.Indeed, the history of science includes many examples of theories that were rejected because they lacked parsimony and/or scope, even before their falsity (and the truth of their competing theory) was confirmed by observation. As I point out in my book, geocentrism was rejected by the scientific community because it lacked parsimony, almost 200 years before the observation of parallax confirmed that the Earth moves around the sun. And the continuum theory of matter was completely rejected by 1860, because it couldn’t explain chemical reactions or organize matter into a periodic table, even though the experiments that eventually confirmed atomic theory weren’t even proposed until the 1870s. Solving the best of a bad lot problemIn response, the “empiricists” are likely to object, this time pointing to the “best of a bad lot” problem. “We have no reason to conclude that an explanation that has been shown to be better than others is true because we have no guarantee that the true theory was among those that we were considering.” If you have a “lot” (a collection) of poor explanations, even the best one is not true.The responses to this problem, however, are numerous. Peter Lipton argues that the set of hypotheses that scientists consider are not just randomly selected and thus unlikely to contain a theory that is true; they use existing knowledge to home in on what is most likely to be true. A doctor trying to explain abdominal pain is not going to bother wondering whether demons are causing it; they are going to stick with things like appendicitis and kidney stones, because they are most likely. Further, despite its name, inference to the best explanation is not just the process of comparing hypotheses and accepting the best one; we also want one that is a good explanation in its own right. If it’s not, we are going to keep looking until we find an explanation that is. And while that process does not guarantee that we will come across the exact truth, it’s very likely that the theory that emerges as the best will be at least approximately true, and thus we will be justified in believing it. And lastly, as I point out in my book, although many theories not based in scientific evidence (like geocentrism and continuum theory) have been completely overturned, science almost always progresses by the refinement of scientifically established theories—not their replacement. (Although our sun is not the center of the universe, heliocentrism was our first step in understanding the Earth’s true place in the universe.) So even if a theory we accept today may not be completely right, accepting it is a step on the journey toward the complete truth. The example of atomic theoryNow, it’s worth pointing out that most empiricists are not skeptical of all scientific theories; they are only skeptical of the ones that hypothesize the existence of “unobservables.” It’s not like they think that round Earth, heliocentrism, and germ theory are just useful (but untrue) “instruments” for making predictions about how to most efficiently navigate the seas, predict planetary motion, and treat diseases. They admit that the Earth actually is round, that the planets revolve around the sun, and that germs exist and cause disease. These facts and entities can be observed; the theories that suggest they exist are true. It’s only when theories start hypothesizing the existence of unobservables—like electrons, or quantum fields—that can’t be observed directly, that we should start doubting the theories’ truth. And to be fair to van Fraassen, in this published work, he doesn’t really say that we should doubt the truth of such theories. Only that scientific reasoning does not rationally force us to accept them as true. Skepticism, or agnosticism, he argues, is warranted. SUGGESTED READING Science is based in metaphor By Andrew Reynolds But the history of atomic theory throws a monkey wrench into that argument. Atoms used to be unobservable; again, scientists first accepted the theory that they exist because it had such grandiose explanatory power. But then atoms were seen with electron microscopes. They thus moved into the category of the observable. But this didn’t substantially change anything about what scientists were justified in believing; it certainly didn’t change any “atomic skeptics’” minds. It just confirmed what everyone already knew! If the hypothesis that something exists clearly provides the best explanation for what we can observe, then the belief that it exists is what is rational. The fact that the thing itself is not technically observable does not make doubting its existence rational.What’s more, the atoms were seen with an electron microscope. But electrons are themselves unobservable. How can an electron microscope reveal an atom if electrons are not real? This leads us to what is famously known as Hilary Putnam’s “miracle argument.” It’s (not) a miraclePhilosopher Hilary Putnam’s miracle argument is somewhat famous. Basically, he argued that a scientific theory being able to repeatedly make successful predictions—even about unobservables—without being at least approximately true would require a “miracle.” In other words, it would be such an improbable occurrence that it could not rationally be believed—the better explanation for why scientific theories about unobservables make the right predictions is because the unobservable entities and structures they hypothesize exist.Van Fraassen has addressed this argument. Indeed, in his original article, where he references Alison Gopnik’s point about orgasms not being “necessary for procreation” even though the desire for orgasm “drives the creation of progeny,” he is pointing to his own argument about how organisms are not naturally selected by evolution to get true beliefs; to the extent that belief-forming processes are selected for, the ones that are selected are the ones that are useful for survival—not the ones that produce true beliefs. In the same way, van Fraassen argues, scientific theories are accepted, not because they are true, but because they aren’t falsified; they make successful predictions; that’s what makes them “stick around.” So, the fact that they make successful predictions (and thus became accepted) is not a reason to think they are true.This argument is similar to C. S. Lewis’s “argument from reason,” which he used to argue against evolution: “The truth of evolution undercuts our ability to know that it is true.” Lewis’ argument famously failed; Elizabeth Anscombe, by Lewis’s own admission, demolished Lewis in a debate on the topic, and I have explained elsewhere why modified versions of the argument fail as well. In short, it’s because it fails to recognize that a belief-forming process can make an organism more likely to survive because it reliably produces true beliefs. It is therefore not the case that beliefs generated by naturally selected belief-forming processes are automatically false or unjustified.In the same way, the reason a theory “sticks around” could be because it’s not falsified, because it “makes correct predictions.” But the reason it makes correct predictions is most likely because it is true. (In pointing this out, I am echoing the arguments of Musgrave, Lipton, Psillos, and Kitcher, who refuted van Fraassen’s argument and analogy long before his 2024 article.) It would be a huge coincidence that it could do so without being true. And that is the point of Putnam’s miracle argument. Thus, contrary to van Fraassen’s suggestion, the mere fact that a theory’s predictive success is why it “sticks around” or “becomes accepted” in the scientific community is not a reason to be skeptical about its truth. ConclusionLet me close by saying something about the fact that empiricists don’t doubt the truth of all scientific theories, just the ones that hypothesize unobservables. In today’s society, this theory is dangerous to espouse publicly. The nuance between being skeptical about all of science, and just the highly technical theories about unobservables, is guaranteed to be lost to the average reader. And there is so much unjustified skepticism and doubt about the findings of science and experts that a headline like “Science Does Not Describe Reality” could easily fuel the anti-science movement, worsen climate change denial and cause the next pandemic.This is why I believe it was important to still respond to van Fraassen’s article, a full two years after he wrote it. Not only, as I have argued, does it represent a philosophically inaccurate view of how scientific reasoning works and what scientific reasoning can do, but it is a dangerous view to boot. It should thus be rejected.

iai.tv logo
Oct 1 • 8:00 AM EDT • Science • iai.tv
Reality Check: “Fifteen” attack ad on Ford raising taxes

The Better Nevada PAC’s assertion that Ford supported 15 tax hikes totaling over $3 billion is largely inaccurate. Several alleged tax increases were actually extensions of previous legislation or involved

2news.com logo
Sep 24 • 10:15 PM EDT • Politics • 2news.com
Why politicians should pay more attention to the polls (reprise)

Welcome to the 231st edition of The Week in Polls (TWIP), a polling newsletter that always welcomes new readers. So please don’t keep TWIP a secret; share it with your friends and colleagues: Share The Week in Polls With all the stories swirling around Reform, JL Partners and political polling, I am finally tearing myself away from talking about the regulation of polling this week - and instead re-running a previous piece reminding us why the solution to problems with polls isn’t simply to ban or ignore them. Rather, we need them to play a vital role in making our politics - and politicians - better. This is followed by summaries of the latest national voting intention polls, seat projections from MRPs and the most recent party leader ratings. Subscribers who pay for the full edition can also read ten insights from the latest polls and analyses, including: A trio of organisations and a pair of professors have been casting their eyes over the Liberal Democrats ahead of the party’s autumn conference; What impact did the summer heatwave have on views on climate change? and Which wins out between chocolate and cheese? If you are not yet a paid subscriber, you can read all ten by starting a free trial now: Get FREE 7-day trial For more polling news as it comes in ahead of next week’s edition, you can also follow The Week in Polls on Bluesky. Picking up on my doubt last week about whether or not 35% of British train travellers have tried to order a takeaway for delivery to a railway station or carriage, thank you to reader Iain, who has been in touch to say he has done just that… in India. Another quick pair of updates, this time on the controversies around Reform and polling. Politico has a further scoop - “Farage aide caught in donor sting attended Egypt retreat with JL Partners” - and Ian Dunt has written about the problems with polling too: “Seeing the response of many polling experts in the last few weeks suggests that they are nervous about what is happening and keen to mitigate against it. Seeing the response of journalists, which has been basically non-existent, suggests the precise opposite.” My usual reminder as well: please do hit reply if you’ve spotted a poll, focus group or new piece of poll-based research that you think should be in contention for coverage in a future edition. Thank you! And with that, on with the show. How politicians get public opinion wrong I have written before about the benefits that come from widespread availability of high-quality political polling. We need more polls and more good coverage of polls. However, what I’ve not dug into before is the research into how accurately politicians estimate public opinion on issues. Thankfully we have academic research to tell us. First, the bad news: We present results from a study of 866 politicians in four countries.1 Politicians were asked to estimate the percentage of public support for various policy proposals. Comparing more than 10,000 estimations with actual levels of public support, we conclude that politicians are quite inaccurate estimators of people’s preferences. They make large errors and even regularly misperceive what a majority of the voters wants. Politicians are hardly better at estimating public preferences than ordinary citizens. Walgrave, S., Jansen, A., Sevenans, J., Soontjens, K., Pilet, J.B., Brack, N., Varone, F., Helfer, L., Vliegenthart, R., van der Meer, T., Breunig, C., Bailer, S., Sheffer, L. & Loewen, P.J. (2023a), “Inaccurate Politicians: Elected Representatives’ Estimations of Public Opinion in Four Countries”, The Journal of Politics, 85(1): 209-222. The bad news is not quite as bad as that makes it sound because, when asked whether the majority of public opinion was for or against various policies, politicians picked the wrong option ‘only’ 29% of the time. However, this included getting it wrong even when public opinion is far from evenly split on an issue: They do not only get it wrong when the public is divided and hard to read (e.g., when there is a 51%:49% distribution). Seventy percent of the misplacements occur for policies with a relatively clear distribution of at least 60% (dis)agreeing citizens. Moreover, the public scored only slightly worse than politicians, picking the wrong option 2.7 out of 8 times on average, compared with 2.3 times for politicians. Turning to the other measure of accuracy (not which side is in the lead but what the percentage support is for each side), politicians on average were 18 percentage points out in their guesses. Again, this was only slightly better than the public, whose error came in at 21 percentage points.2 Further research, using the same mix of countries, gives a little more comfort about how (in)accurate perceptions are: We show that politicians have a better understanding of public opinion when they think the issue matters to voters. Further, when an issue is personally important to politicians they more accurately estimate their party supporters’ opinions. The results confirm that politicians hold more accurate perceptions of voters’ preferences when they think it is important to do so but not necessarily when the issues actually are important to voters. Butler, C., Walgrave, S., Soontjens, K. & Loewen, P.J. (2024), “Politicians are better at estimating public opinion when they think it is more salient”, Party Politics: 1–13. Even so, it’s not a good picture. The remedy for this? Politicians should spend more time reading the polls. So forward this newsletter to someone you know… Share For a more general argument against banning polls, see my 2024 piece. Is “Let Bartlet Be Bartlet” good advice or not? The latest episode of my Political Fictions podcast heads across the Atlantic for a dip into one of the political TV classics: Mark and Cory talk about a giant of American political drama for the first time this week, as they tackle The West Wing. They talk about Let Bartlet Be Bartlet, the nineteenth episode of series one. In it, Mark wonders if The West Wing is actually the epitome of the “Boris Johnson school of politics”, Cory outmanoeuvres Mark into confirming we will record a Christmas Special about Love, Actually and they debate whether the central message of this episode has been misunderstood. Cory’s email newsletter is Paperback Rioter and Mark has a family of email newsletters. Subscribe to our newsletter for some additional bonus content and to get alerted whenever a new episode comes out. Our theme tune is “Monkeys Spinning Monkeys” by Kevin MacLeod/incompetech.com and licensed under the Creative Commons: By Attribution 4.0 License. You can also listen to our episode on Apple Podcasts, Spotify, the web or YouTube. Elsewhere from me… [ A Lord's Eye View How to fix the problem of money in our politics Welcome to my latest update on work in the House of Lords. This time I am taking a look at how our political system is regulated, and the protections we need to keep democracy safe from becoming simply a contest to see who has the biggest bank account… Read more 5 days ago · 3 likes · 2 comments · Mark Pack ](https://lordseyeview.substack.com/p/how-to-fix-the-problem-of-money-in?utm_source=substack&utm_campaign=post_embed&utm_medium=web&embedding_publication_id=1269283) If you spot a factual error… Borrowing an idea from Stuart Ritchie and others, if you spot a factual error in an edition of this newsletter, I will give you a free 12-month subscription to the paid-for version (or an extra 12 months on your existing one). Factual errors qualify, but grammatical errors or disagreements over interpretation do not, and my decision is final. Messages about both of those are always welcome, though. Voting intention and leadership ratings Here are the latest voting intention polls from different pollsters. It is all pretty static at the moment, with a small and consistent Labour lead in nearly all the polls. The table is also online here, along with explanations about what is included. It is updated regularly throughout the week as new polls are released. Next are the seat projections from MRPs, also sorted by fieldwork dates, listing those which are based on dedicated fieldwork (rather than models based on averages of other polls). Because such MRPs are published infrequently, some of the entries below are quite old. Finally, here is a summary of leadership ratings, with a graph from PollCheck: Detailed numbers from the latest polls are online here, and both that and the graph are updated regularly throughout the week as new data comes out. For historical figures on both voting intention and leader approval stretching back to the 1930s, including Parliamentary by-election polls, see PollBase. Catch up on the previous two editions [ ](https://theweekinpolls.substack.com/p/a-crisis-point-for-political-polling) [ A crisis point for political polling Sep 13 ](https://theweekinpolls.substack.com/p/a-crisis-point-for-political-polling) [ Read full story ](https://theweekinpolls.substack.com/p/a-crisis-point-for-political-polling) [ ](https://theweekinpolls.substack.com/p/extra-questions-about-the-polls-in) [ Extra questions about the polls in Reform’s funding scandal Sep 6 ](https://theweekinpolls.substack.com/p/extra-questions-about-the-polls-in) [ Read full story ](https://theweekinpolls.substack.com/p/extra-questions-about-the-polls-in) My privacy policy and other legal information are available here. Links to purchase books or other items are usually affiliate links which pay a commission for each sale. Quotes from social media messages are occasionally lightly edited for punctuation and clarity. If you are subscribed to other email lists of mine, please note that unsubscribing from this one will not automatically remove you from the others. If you wish to be removed from all my lists, simply reply to this email to let me know. A quintet of Lib Dem polling analyses, and 9 other polling insights you shouldn’t miss The following ten findings from the most recent polls and analyses are only for subscribers who pay for the full edition, but you can sign up for a free trial to read them straight away. Ahead of the start of Liberal Democrat conference, we have a trio of organisations and a pair of professors with pieces out about Ed Davey’s party:

theweekinpolls.substack.com logo
Sep 20 • 1:02 PM EDT • Politics • theweekinpolls.substack.com
The Big Holes in Wearable Heart Rate Variability And Readiness Scores

On September 9th, Apple announced it was revamping its Apple Watch Health Sensing System, rolling out a Readiness score (0 to 10), and increasing the frequency of heart rate variability (HRV) outputs 24-fold. This can be viewed as upping its competition with various other consumer wearable sensors. These “readiness scores” as a composite of multiple metrics, with heart rate variability (HRV) being front and center for most. The majority of Americans are now using wearable sensors, which equates to well over 100 million adults. HRV and Readiness scores are increasingly being marketed as a measurement of autonomic nervous system health, a digital marker for future disease, a clock for biological age, and a holistic metric to promote healthspan and even longevity. (Apple also introduced a new longevity tab and “Health Age”.) None of this has been proven. In this edition of Ground Truths I am going to review what we know about heart rate variability and readiness scores. Heart Rate Variability HRV is the variation in normal heart cycle timing. The variability of the heart rate, the barely perceptible millisecond changes in time between consecutive heart beats (see R-R intervals in the Figure below, left panel), is due to interplay between the sympathetic and parasympathetic (vagal nerve) inputs. Distinct from heart rate, individuals with the same heart rate can have very different HRVs. It is a rough reflection of the autonomic nervous system (ANS) activity, inadequate to say whether a person’s ANS function is abnormal. For more than three decades, heart rate variability (HRV) has been measured and several studies have found an association of low HRV and clinical outcomes, particularly a link with higher all-cause and cardiovascular mortality. There have also been less well established links of low HRV to risk of early cognitive impairment, dementia, mental illness, Type 2 diabetes and substance abuse. An important reminder is that HRV is a surrogate marker without any established cause-and-effect relationship. If you increase your HRV, that doesn’t mean it will improve health outcomes. In fact, there is no hard evidence for that. All that work linking to health outcomes was done with electrocardiogram (ECG) derived HRV. Now, in the era of consumer wearables, this is getting assessed differently, by optical pulse (yes, the lights you see) plethysmography (PPG) or what is called pulse rate variability (PRV). They are not the same, as shown below (right panel) and only concordant when the delay between the ECG and pulse is kept constant, which basically means at rest. I should mention there’s also what I will call MPV, a mattress mechanical movement sensor, a derived heart rate variability, that companies like Eight Sleep use, even further away from directly measuring HRV. HRV has not one uniform measurement but many different types of quantification, such as RMSSD, the magnitude of difference between successive R-R intervals of normal sinus beats (N-N) or SDNN, the standard deviation of NN intervals, both in milliseconds. SDNN is one of the so-called frequency domain HRVs (others are LF, HF, LF/HF). Different wearable sensors use different metics; Apple has relied on SDNN and nearly all of the others use RMSSD, which is generally considered the more accurate metric. There’s also the different length of time measured, such as for a matter of minutes, all day, or an overnight’s sleep. Short measurements are especially problematic since they don’t capture enough of respiratory modulation and other factors that influence HRV. Share Ground Truths How well does HRV correlate with PRV? There are very limited studies, especially independently done. One that is commonly cited was conducted by Air Force researchers in only 13 healthy adults assessing Oura ring 3 and 4, Whoop 4.0, and Garmin Fenix 6 and showed a correlation coefficient of 0.88 to 0.97 and a mean absolute percentage error from 6 to 10%. The correlation is not a perfect 1.0, but there’s at least a fairly high level of correlation. HRV is supposed to increase during the night due to takeover of the parasympathetic nervous system, and higher during deep sleep. A recent example of my 1 week, all day “HRV,” and one during sleep is shown below. As you can see, the N of 1 data are inconsistent for the same days from different sensors (Oura, AppleWatch, Fitbit Air, and Eight Sleep) by patterns, absolute numbers, and comparison with prior days and weeks. The largest study in over 8 million Fitbit users (the old version, not Google Fitbit Air, introduced in May 2026) gives you a sense of the effect of age, sex, and the 2 different main HRV (here PRV) metrics, with RMSSD on the left and SDRR (=SDNN) on the right below. That study, from data collected in 2018, is a major outlier, since all the more recent ones are tiny with respect to sample size. Many of the companies have not had independent evaluation of their HRV, such as Eight Sleep, but have published a low standard error on their website. There are some other published studies on the correlation between HRV and PRV, but they are all small and only in healthy adults. A scoping review emphasized the lack of study in underrepresented individuals, including the aged, people of color (which affects the PPG signal), and individuals who are underweight or obese. Add the typical adult age 60 plus with one or more chronic diseases. For example, one study in over 900 adults found poor correlation of HRV and PRV, non-uniformly underestimated across many chronic diseases (cardiovascular, endocrine, neurological, respiratory, and others), concluding PRV is “an invalid surrogate for HRV.” A recent systematic review of 43 studies comparing HRV and PRV found reasonable pooled absolute standardized error (HRV as gold standard) but only 10 of the studies provided quantitative synthesis in ideal conditions. Their main conclusion was similarly cautious: “PPG-derived HRV [PRV] should not be regarded as universally interchangeable with ECG-derived HRV across all devices, populations, and recording contexts.” Factors Affecting HRV and PRV That gets me to the long list of factors that affect HRV (and PRV) besides the device, the type of measurement (RMSSD, SDNN or others), the person’s signal, the sensor site, the duration of data capture, if weighting by sleep stage is used, how artifact is processed and corrected. And this list is not complete!: Oura puts out data from their community of users (who input data) on what affects their overnight HRV. The factors currently provided are: no alcohol (increase 8%), melatonin (increase 2%, float tank (increase 2%), wine (decrease 4%), and party (decrease 14%) in overnight HRV. Must be some big parties! Subscribe What is a PRV measurement good for? It has been falsely characterized as an index of “autonomic balance” and a specific indicator of stress. A 2018 review of the studies available for HRV and its relationship to stress, not using any of the current wearables, found that stress can lower HRV. But so can many other factors. The non-specificity of the signal, indexed to the table above, is striking. Evidence from a UK Biobank study of over 46,000 participants with actual HRV looked at genetically predicted HRV, a genetic risk score, that failed to show the expected HRV-mortality link, indicating that _HRV is likely not causa_l, but rather a reflection of person’s physiologic state. A review of consumer wearable HRV data from 5 longitudinal studies showed that nighttime PRV was not associated with perceived stress, and surprisingly higher HRV, in the largest cohort (N=717 participants), was correlated with higher stress. An Oura ring cohort of 525 first-year college students found a link between overnight PRV and perceived stress, but that was also seen with resting heart rate, sleep, and respiratory rate. Several very small studies have examined the relationship of HRV and athletic injuries or guiding training with mixed, and predominantly negative results. HRV biofeedback training with paced breathing had no significant effect on reducing stress or raising HRV, as demonstrated with sham controlled trials. When HRV for multiple days showed a decline in conjunction with body temperature, the Oura ring published data for prediction of Covid. The WHOOP company sponsored an observational study, published in 2026, of 30,000 users for 72 weeks, without a control group, that reported reduced alcohol intake (5.8 % points) by self-report. That doesn’t tell us much, and particularly about the merits of HRV for behavioral change. If you use the same device and conditions as longitudinal trends for multiple (at least 2-3) weeks that may be the one way to get something useful from the measurements. Data for overnight sleep with minimal motion and using RMSSD is the best proxy for real HRV. The reason to look at trends rather than any given night is that it more likely represents something, even though you won’t know with certainty what the “it” is. Keep in mind there are no data, no peer-reviewed evidence, to show that HRV fluctuation in-person has any correlation with health outcomes. Share Readiness Scores These are proprietary scores that integrate different metrics for each of the wearables: no algorithms have been disclosed. They are unvalidated against health outcomes. In a review of 14 composite health scores of readiness and recovery, HRV contributed 86% to the scores, followed by reading heart rate (79%), physical activity and sleep duration (both at 71%). That review noted the substantial variability in measurement protools and lack of standardization. Sleep staging is notoriously inconsistent and inaccurate by these sensors, which adds further to the HRV uncertainties for what the scores, which use sleep stage data, mean. Only resting heart rate has been shown consistently across devices to be extremely accurate. I’ve made a Table to summarize what we know about which metrics are included, the scores, any peer-reviewed studies that compared the readiness score with health outcomes, and the corresponding (if any, NA-not available) citation. You will note that some companies do not use the term “readiness," such as WHOOP for recovery, and Garmin, which has 2 different scores, one of which is Body Battery. Eight Sleep uses the term “Fitness Score.” They all include HRV; Apple includes a new metric they call “Recovery HRV” which among other components uses 7-days of sleep, but it is unclear what this means or how it is differentiated from other scores (there are clearly no data for outcomes). We have no knowledge of how the different components are weighted or whether any of these scores are better than resting heart rate, HRV alone, physical activity, or any other single metric. Since none of these are standardized, they are not interchangeable, so if you get a 90 for Oura that has no relationship to a 90 on a Google Fitbit Air. Notably, the company can update its algorithm for readiness score at any point without notification to device users. Without any useful evidence of actionability for these scores or established relationship with health outcomes, it is hard to make a case for their value. At the Apple recent announcement they showed their Readiness score (0-10) on the watch (Figure below) but there are no published data on this score, not even on their website. It’s available only on their new Watch Series 12 or Ultra 4 [of course, ;-)]. That exemplifies the problems with these scores, lack of data and evidence for being meaningful to promote health. Perhaps the best study (which isn’t saying much) is the WHOOP Recovery for golfer performance, because it did correlate with an objective outcome, even though there was no control group and the authors were all from the company. Among the 389 pro golfers, an absolute 10-per cent point increase in Recovery score was associated with about 0.5 fewer strokes per round. But that’s hardly a health outcome! WHOOP is also conducting a study in over 2,700 runners to see if their recovery score will be linked to less injuries and improved performance, but that is not yet published and has no control group or randomization. Putting This in Context For two decades I’ve been enthusiastic about the potential for digital health and particularly wearable biosensors. Over the years, we’ve seen some great progress for their ability to promote physical activity and accurately detect atrial fibrillation (the first FDA cleared deep learning AI for consumers). That work was the subject of rigorous research. But there are holes in the data and evidence for other metrics. One notable one is the “VO2 max” story that I wrote about earlier this year. At that time many subscribers asked me to cover heart rate variability, which I finally got to here. When I dived into the research and publication for HRV and readiness scores, I expected to find at least some that were of high quality and demonstrated their utility by linkage to health outcomes. To my surprise, I found none. The wearable sensor measurements for HRV (PRV) are, for the most part, accurate, but that validation work has only been done in small studies of healthy adults and does not take into account the long list of factors, from the device, software side, and the user side, that affect HRV measurements. Moreover, this metric chiefly relies on optical sensing and, as we have learned for heart rate PPG sensing, may be less accurate in people of color. Keep in mind that all of the health outcome association evidence comes from ECG-derived HRV; none are from wearable sensor data. I will repeat the key point: there's no peer-reviewed evidence to show that in-person HRV fluctuation—or efforts to raise your HRV— has any correlation with health outcomes. For those of you who look at your HRV on awakening, or even 2+ week trends of it being low, I hope this context helps to relieve any anxiety. Yes, low HRV (not PRV) has been shown to increase risk of some diseases as summarized above. But efforts to raise your HRV—a surrogate metric— has not been established for improving any health outcomes. HRV does not have any evidence of causality (the genetic evidence actually goes against this possibility). In this summary, I have not included data for other wearable sensors such as Polar, Samsung, Withings, Suunto, Amazfit, Coros, Ultrahuman, or additional mattress sensors. These are beyond my first-hand experience, and as far as I know from my in-depth review none have any peer-reviewed published data that differ from the 6 sensors I’ve reviewed here. From the points I’ve gone over above, we’re not ready for readiness scores. Besides being proprietary, they are predominantly based on metrics that have their own issues. It’s compounding the problem, like building a house without a solid foundation that has never undergone a rigorous inspection, and then selling it. Like I mentioned for PRV, you can look at trends over weeks rather than any single day, to get a handle, but even that may not be helpful. There’s simply no evidence that these scores meaningfully relate to health outcomes. I’d emphasize they might, but that requires doing prospective or randomized studies to prove it. None exist. There’s great promise for HRV/PRV utility_. For example**,**_ Prof Maiken Nedergaard, who discovered the brain glymphatics that are essential in eliminating metabolic waste products from the brain during sleep, has posited that HRV could be a non-invasive marker for neuromodulator oscillations, brain-body regulatory circuits, and brain clearance. That would be extremely useful, but like everything else on HRV and readiness scores it requires solid research and validation. The lay media isn’t helping much to get the story straight. Earlier this year The Economist published a piece entitled “The most useful indicator of your overall health” which ordained HRV as an “accumulated stress score.” That’s akin to the false assertion about VO2max: “V02 max is the singular most powerful marker for longevity.” As I’ve summarized here, that is not established. The fact is that so many things can lower HRV, including physical exercise (especially an intense workout), reduced sleep quality, stress, the list above, no less the device, signal, and software. Whatever fluctuations observed have not been correlated with any health outcome. Sadly, “datamaxxers” are widely using HRV and readiness scores that have never been validated to mean anything. We already know that for some people using the sensors for sleep metrics, it can induce “orthosomnia,” an obsession to get high sleep quality, with associated high levels of anxiety. In an experiment done by a company to promote sleep quality for its employees, “For those employees who did use the trackers, many reported feeling perfectly rested until their tracker told them they had had a terrible night. Others were told that they had slept like a baby when they had actually been lying awake worrying about the quality of their sleep. “ The same problem can result from preoccupation with HRV or readiness scores, with anxiety that would lead to further reduction in both. It you are using a wearable like >100 million American adults, it’s OK to look at these data, but contextualized with the major caveats reviewed here. If you are one to require evidence that HRV or readiness scores are linked to health outcomes, you may not even want to look. The companies make it hard to turn them off! Let me end with the companies that make and sell wearables. Apple’s doubling down on HRV (24-fold more reporting and heart rate very 5 seconds) and introduction of a Readiness score tells us that consumers have bought into these metrics and they are joining the club. However, all of this is occurring with a backdrop of tens millions of users, claims about the data that are not backed up by adequate evidence, marketing way out in front of whatever limited data exists, and not being transparent about their readiness score algorithms. The companies can well afford to do the research that is needed to connect these metrics with health outcomes show, once and for all, that increasing HRV or using readiness scores promotes our health. If they believed and invested in the products they are selling, we’d not be in this position of not knowing. That’s essentially where we are with HRV and readiness scores. Perhaps someday this will change and we’ll have good reason to embrace them. NB: I wrote this post. No AI. I have no conflicts of interest with any of its content. Loading... Ground Truths has 215,000 subscribers from every US state and 214 countries. There are over 300,000 followers of Ground Truths so more than 90,000 folks who can easily convert to be free subscribers. Your subscription to these free essays and podcasts makes my work in putting them together worthwhile. If you’re not a subscriber, please join! If you found this interesting PLEASE share it! Share Ground Truths The proceeds from all voluntary paid subscriptions go to support our summer internship program. It enabled us to accept and support a record number of 62 summer interns that joined us in 2026! These are high school, college and medical students selected from thousands of applicants. We couldn’t do this expanded program without the funds coming in through Ground Truths. Thank you!

erictopol.substack.com logo
Sep 19 • 10:05 AM EDT • Technology • erictopol.substack.com
Fact-Checking Trump’s Attacks on Democratic Senate Candidates

The president attacked three Democrats running for Senate this week with inaccurate claims to portray them as too extreme.

nytimes.com logo
Sep 11 • 5:12 PM EDT • Politics • nytimes.com
ABC denies 'inaccurate' report 'Jimmy Kimmel Live!' will end in 2027

Amid speculation about Jimmy Kimmel's late-night tenure coming to an end, ABC is pushing back. On Sept. 8, Page Six reported that this season of "Jimmy Kimmel Live," will be

centraloregondaily.com logo
Sep 8 • 6:51 PM EDT • Entertainment • centraloregondaily.com
Different State Regulatory Approaches Reflect Open Questions About AI Mental Health Tools

States are taking different regulatory approaches to AI mental health tools in response to concerns about chatbots providing inaccurate or potentially dangerous advice. Laws and pending legislation in some states restrict AI from providing or advertising itself as therapy, while others focus on data protections, disclosures, and patient consent requirements.

kff.org logo
Aug 27 • 8:00 AM EDT • Health • kff.org
Google's AI Overviews are causing chaos for small businesses

Google AI Overviews are becoming customers' first impression of small businesses. Inaccurate summaries are damaging their reputations.

businessinsider.com logo
Aug 11 • 4:22 AM EDT • Business • businessinsider.com
🙂 Positive (50%)
How vaccine skeptics use storytelling to create powerful and medically inaccurate messaging

Recognizing the stories behind people’s concerns with vaccines can help public health advocates build a more inclusive approach to healthcare.

washingtonpost.com logo
Jul 27 • 10:09 PM EDT • Health • washingtonpost.com
Trump orders signs placed outside Smithsonian museum warning of "inaccurate" information
cbsnews.com logo
Jul 24 • 7:38 PM EDT • Politics • cbsnews.com
High-Signal Publisher
Ohio business says inaccurate Google AI summary of store is hurting sales

“The fact that we know about it, we are fighting it and the community is rallying behind us is the reason that we're able to continue to be in business,” Chanel Bush, who opened the Viridian Bookshop in October of 2025, said.

spectrumnews1.com logo
Jul 23 • 5:38 PM EDT • Business • spectrumnews1.com
FDA is still focused on lettuce supplier as source of parasite, despite faulty test result

Federal health officials said Monday they remain focused on lettuce from Taylor Farms as the source of a multistate outbreak of a diarrhea-causing parasite, despite inaccurate test results that the government briefly publicized over the weekend.

pbs.org logo
Jul 20 • 4:18 PM EDT • Health • pbs.org
🙂 Positive (35%)
Shaky political "science": breakdown of the peer review process

[This article is the latest in our “Shaky Political Science” series. In our 38 page report, “Shaky political ‘science’ misses mark on ranked choice voting,” released in December 2025, we examined over 40 studies focused on ranked choice voting (RCV). Many of these studies suffered from puzzling research methodologies, poorly constructed surveys, unrealistic simulated elections, cherry picked data and faulty analyses that often were contradicted by results from real-world elections (a link to our full paper is here). Here are links to the first four articles in our series: the first one summarizing the overall results of our report; the second one analyzed a study co-authored by University of Minnesota academic Larry Jacobs, which was one of the most error-prone of all the research we reviewed in our report; a third article showing an admirable example of two political scientists holding two other political scientists accountable for their flawed, sloppy research; and a fourth article that analyzed seven studies sponsored by New America’s political reform program that deployed flawed methodology based on “simulated” elections instead of real-world elections, which when combined with some basic misunderstandings of RCV, produced odd and misleading results]. In this fifth article in our “Shaky Political Science” series, we examine two flawed studies on the impact of RCV on voter turnout authored by San Francisco State University professor Jason McDaniel. McDaniel’s research in these two papers stands out for its puzzling methodology, inadequate and cherry-picked data, and faulty analyses that not only misunderstands ranked choice voting but reflects a basic misunderstanding of politics in general. The first of the McDaniel studies actually passed through a peer review process and then was subsequently published by the Journal of Urban Affairs, despite its obvious errors. It is perplexing that a jury of peer political scientists did not catch the errors and send it back to McDaniel for a rewrite or outright rejection. This calls into question the credibility of the peer review process itself. The peer review process is supposed to represent the intellectual pinnacle of the profession – a way for academics to certify the quality of each other’s work, and to act as a check against sloppy research. In the case of the McDaniel study, the peer review process failed. Yet his flawed study has been cited dozens of times by other political scientists in their own flawed works — an unvirtuous circle of academic research. Here is the first Jason McDaniel study that we analyzed: “Writing the Rules to Rank the Candidates: Examining the Impact of Instant-Runoff Voting on Racial Group Turnout in San Francisco Mayoral Elections,” by Jason McDaniel, 2016, Journal of Urban Affairs. This study covered San Francisco’s use of ranked choice voting (RCV) in mayoral elections, purporting to show that RCV reduces voter turnout overall, increases voter errors, and has a negative impact on “marginal populations,” i.e. racial minorities, mainly as the result of the alleged complexity of RCV asking voters to rank candidates and having to be familiar with multiple candidates and related dynamics. McDaniel reached these conclusions by examining five San Francisco mayoral elections taking place in odd years (i.e. where there are no federal or state races on the ballot) from 1995 to 2011, comparing three earlier non-RCV races (in 1995, 1999 and 2003) with the later two RCV races (in 2007 and 2011). Only two RCV races were studied, and as McDaniel acknowledged, only the contest in 2011 was “competitive” (that contest in 2007 was the first mayoral election using RCV, and incumbent mayor Gavin Newsom won easily in the first round of counting, earning 74% of first choices and far ahead of the second place candidate who had just 6.3% of first choices). Yet even the allegedly competitive contest in 2011was won by an incumbent with a landslide margin of nearly 20 points. No other elections were studied by McDaniel, even though San Francisco elects 17 other offices besides mayor with RCV, including the 11 seats on the Board of Supervisors (the name for the city council) in even years, and six other citywide offices in odd years. Previous academic research has found that competitive elections can stimulate greater voter involvement and higher turnout. For unexplained reasons, this study ignored at least 34 other RCV races between the years 2004 to 2011, many of which were very competitive and had high voter turnout. Instead, McDaniel’s analysis based its conclusions on data from only a single, marginally-competitive mayor’s race using RCV. These 2007 and 2011 mayoral elections were the only two RCV elections that McDaniel used as the sole basis for his small data set to determine the impact of RCV on voter turnout. Despite having no meaningful data to inform his study, McDaniel nonetheless concluded that RCV lowered turnout when compared to the previous three mayoral non-RCV elections using two-round runoffs. Then, using the same inadequate database of San Francisco elections, McDaniel also found that “Instant-runoff voting is associated with a significant decline in Black voter turnout and White voter turnout…On average, Black voter turnout decreases by 18 points and White voter turnout decreases by 16 points …” But how could he credibly arrive at that conclusion when his study was based on voter turnout in two RCV contests that were not even remotely competitive? Beyond that, McDaniel tacitly admits that the key determining factor driving Black voter turnout was not RCV vs non-RCV elections, but the presence of a viable Black candidate. McDaniel wrote, “Black voter turnout was about 17 percentage points higher for the two elections that featured an African American candidate compared to the three elections that did not” — but one of those three elections did not use RCV. The same dynamic was present for White voter turnout, in which higher turnout was not correlated with RCV vs non-RCV elections, but the presence of a viable White candidate. Beyond that, McDaniel’s voter turnout results, whether correlated with the race of the candidates or RCV vs non-RCV elections, showed no clear patterns. In his Figure 1, he shows low Asian voter turnout in the non-RCV elections in 1999 and 2003 when there was no viable Asian candidate running, but in the 2007 mayoral election using RCV, there is still no viable Asian candidate, yet Asian turnout increases dramatically. For Latinos, voter turnout is virtually the same in 1999 (no viable Latino candidate, non-RCV elections), 2003 (leading Latino candidate, non-RCV elections), 2007 (no viable Latino candidate, RCV elections) and 2011 (leading Latino candidate, RCV elections). Meanwhile White voter turnout increased dramatically from a low in the 2007 mayoral election, in which the leading candidate was a white incumbent, to the 2011 election in which there were no leading white candidates that finished any higher than sixth place with 5.6% of the vote. Clearly other factors were determining voter turnout than either the race of the candidates or RCV vs non-RCV elections. Despite these inconsistent findings, McDaniel boldly marches forth to render conclusions about the impact of RCV on different groups of racial voters, even though his conclusions are not remotely supported by his data. This looks like the worst kind of “data cherry picking” to cover up a puzzling conclusion based on a poor data set involving only two RCV elections, neither of which were remotely competitive. McDaniel’s cherry picking didn’t stop there. In calculating voter turnout for the three non-RCV mayoral elections in 1995, 1999 and 2003, he failed to mention that in each of these years there was a two-round runoff election, with a first round election in November in which no candidate garnered a winning majority of the vote, followed by the decisive runoff in December. So there were two non-RCV elections in each election year. But McDaniel only decided to include one of the two elections in his data set, apparently the December runoff in each year (which had higher turnout in two out of the three mayor’s races, and lower turnout in the third). So he ignored the November 1995 election, the November 1999 election when turnout was only 45 percent, and the November 2003 election when turnout was only 45.7 percent. This is flawed methodology based on blatant cherry picking, since McDaniel simply omitted three elections worth of data that did not fit within his frame. His study should never have made it through a peer review process, and it is an embarrassment to political science that not only was it approved for publication, but it has been uncritically cited by dozens of other political scientists. Share A “study” without wider context or credible data Lacking a data set that included even a single competitive RCV mayoral race, McDaniel declared that RCV lowered voter turnout in 2011, when turnout was 42.5 percent. But he also ignored -- or did not realize -- that voter turnout in mayoral elections during 2011 had sharply declined in all big cities, including far more steeply in nearby Los Angeles and other cities. In fact, San Francisco’s election in 2011 had the second highest turnout of any mayoral election in the nation’s 22 largest cities from 2008-2011. Was 2011 just a particularly low turnout year all across the country, perhaps because the country was still reeling from the aftermath of the housing market collapse in 2008-2010? McDaniel ignored this factor and did not search for alternative explanations beyond his flawed thesis. Furthermore, while McDaniel suggested that RCV leads to more “overvotes” – which is an invalidated ballot in which a voter, in either a non-RCV or RCV race, makes a mistake and selects two candidates for a single choice – he failed to address that in the June 2012 US Senate primary contest, which used the typical “one-choice” plurality method and did not use RCV, there was more than five times as many overvotes in San Francisco and Oakland than in the RCV mayoral elections in those cities. Additionally, turnout had been significantly higher and more representative in RCV elections for the San Francisco Board of Supervisors races -- that McDaniel didn’t even bother to study -- as compared to previous “delayed two-round runoffs” for those offices, often higher by as much as 40% in the RCV races. Perhaps the most surprising irony was that an earlier study co-authored by McDaniel (Overvoting and the Equality of Voice under Instant-Runoff Voting in San Francisco by Francis Neely and Jason McDaniel, 2015, California Journal of Politics and Policy) found that overvotes are often more common in lower-income precincts, but that the pattern of overvoting is similar in both RCV and non-RCV contests. Wrote Neely and McDaniel: “[Overvote] errors appear to be a function of complexity in general and not IRV per se…. the pattern of overvoting is similar in both IRV and non-IRV contests…Even among votes cast in a top-of-the-ticket race like U.S. Senate, overvoting occurs in an uneven fashion” (IRV, or ‘instant runoff voting,’ is another name for ranked choice voting in single-winner contests). So McDaniel not only contradicted his own conclusion in his earlier study with Neely, but the earlier study had been an important corrective for one of the most frequently made mistakes made by anti-RCV political scientists – a continual failure to provide crucial context by pairing a comparative analysis of real world RCV elections with an analysis of real world plurality or other electoral methods. Unfortunately McDaniel’s earlier study does not get cited much by other researchers, and McDaniel did not even cite it in his own flawed study analyzed above. Incredibly, McDaniel cherrypicked McDaniel, even as his flawed study has been cited dozens of times. Despite being a textbook example of poor political “science,” featuring both inadequate as well as cherrypicked data, the shelf life of McDaniel’s study has been extended endlessly by other political scientists and media outlets who have cited it uncritically. Equally troubling, the Journal of Urban Affairs put this paper through its peer review process and published it despite its numerous flaws. McDaniel gets it wrong, Part II Not content with producing one shoddy, confused paper, McDaniel produced another one a few years later, here is the link: Electoral Rules and Voter Turnout in Mayoral Elections: An Analysis of Ranked-Choice Voting by Jason McDaniel. San Francisco State Univ. (McDaniel, 2019). In this paper, McDaniel extended his previous flawed research focused on San Francisco, about voter turnout and voter impact, to multiple RCV cities through the 2018 elections. In his data set he included San Francisco, Minneapolis, St Paul, Oakland, Berkeley, San Leandro and Santa Fe. He then compared their turnout in mayoral elections to non-RCV cities, and to mayoral elections in those cities held before first RCV usage, and concluded “a significant decrease in voter turnout of approximately 3–5 percentage points in RCV cities after the implementation of RCV.” But he replicated many of the same mistakes as his first study presented above. While in this study McDaniel attempted to distinguish between various factors, such as odd year vs. even year elections, competitive races vs. noncompetitive, contests with open seats and other factors, the analysis is superficial and inadequate. He attempts to incorporate so many often-conflicting variables that his conclusions are once again unsupported by the actual data, and ultimately unconvincing (this second study never underwent a peer review process). For example, McDaniel declared that for RCV cities “in elections that are more competitive than average (margin of victory less than 20%), there is no significant RCV effect on voter turnout.” Yet McDaniel is a professor in a city, San Francisco, in which the 2018 special election using RCV to fill a mayoral vacancy was exceedingly close, decided by a margin of only one percentage point, and resulted in the second highest turnout in San Francisco history for a mayoral election. Illustrating the complexity involved in assessing the critical factors in turnout, that special election was held on the same election day as federal and state offices, which tends to boost turnout. Yet more voters participated in that RCV mayoral election than voted on the same day in the non-RCV primary for governor in which former San Francisco mayor Gavin Newsom was running. Even more puzzling, in this study McDaniel inaccurately cited his own previous study from 2015 with Neely. McDaniel wrote, “Neely and McDaniel (2015) examine individual ballot image files within a natural experiment, and find higher rates of ballot errors in RCV elections compared to non-RCV elections” (emphasis added). However, we quoted from that 2015 study several paragraphs above, and in actual fact Neely and McDaniel actually concluded the opposite -- that the pattern of overvoting – which is the most common form of ballot error – is similar in both RCV and non-RCV contests. McDaniel also ignored the fact that in Bay Area RCV elections held in even years – at the same time as federal and state elections, when turnout is highest -- the “undervote” (that is, the dropoff in ballots cast for governor or mayor) has declined, which contradicts his thesis that the act of ranking ballots in RCV elections is too complex and presents a barrier for certain demographics of voters. To his credit, McDaniel actually acknowledged another contradiction in his own paper, which found a negative RCV turnout effect in some odd-year elections but not others, admitting that “may present a challenge to the theory that increased complexity [from RCV] will have a marginal negative effect on voter turnout.” But he tried to wallpaper over that contradiction by blaming it on an unlikely source -- bad election administration. McDaniels wrote, “This would suggest that the negative impact of RCV on voter participation may be alleviated by high quality election administration,” without presenting any data or even anecdotes to support such an unfounded theory. While competent election administration of any election, whether RCV or non-RCV, is crucially important, there is no credible research that supports the far-fetched notion that the quality of election administration has much of an impact on voter turnout in US elections, whether using RCV or otherwise. Upgrade to a $5 subscription The unvirtuous circle of academic research It is puzzling how McDaniel’s first study could have survived the peer review process for the Journal of Urban Affairs. The mistakes in McDaniel’s first study above should have been obvious to anyone with a basic understanding of ranked choice voting as well as elections in general, and elections in San Francisco specifically. It undermines faith in the credibility of the peer review process to think that a jury of peer political scientists did not catch the egregious errors and reject McDaniel’s sloppy research and paper. And it remains a black mark on political science that this thoroughly discredited study has been cited dozens of times by other researchers. Indeed, in our research paper it became apparent that many political scientists cite each other’s flawed research to pad their list of citations and to legitimize their own flawed research, creating an unvirtuous circle of political science. More broadly, McDaniel’s voter turnout studies, like other studies evaluating RCV’s impact on turnout, lacked a recognition that no credible opinion or research can claim that RCV increases voter turnout merely because so many voters are excited to vote for their favorite candidates instead of “lesser evil” candidates without worrying about spoilers. Instead, with respect to turnout, the advantage of RCV has always been that, in most places where it has been used, it has eliminated a pre-November primary or post-November runoff, which usually saw about half the turnout of the November general election. This is especially true during even years when there are state and federal elections on the ballot which substantially drive turnout more than local contests. Any study that does not adequately account for this non-November primary/runoff vs November general election dynamic, or odd year vs even year election factors, is simply not credible. Both of McDaniel’s studies largely failed on those grounds. Steven Hill and Paul Haughey Share Upgrade to a $5 subscription Thanks for reading DemocracySOS…your digital portal for the pro-democracy movement. Subscribe for only $5 per month to receive full benefits and to support our work. Subscribe

democracysos.substack.com logo
Jul 20 • 9:29 AM EDT • Science • democracysos.substack.com
How Americans are engaged with news, politics, religion and civic life

The Briefing ☀️ Happy Thursday! The Briefing is your guide to the world of news and information. Sign up here! In today’s email: 🔥 Featured story German regulators said this week that the country’s media laws apply to content generated by artificial intelligence after a court ruled last month that Google is liable for inaccurate […]

pewresearch.org logo
Jul 16 • 3:56 PM EDT • Politics • pewresearch.org
Xbox Sends Out Brief Statement Regarding 'Inaccurate' id Tech Reports

They say dozens of people are still working on the engine

purexbox.com logo
Jul 10 • 12:00 PM EDT • Technology • purexbox.com
Anyone can fake a scientific image with AI, tricking even academic journals – and undermining trust in science

Several scientific fields rely on visual evidence to illustrate their claims. Inaccurate AI-generated images put the credibility of science at risk.

theconversation.com logo
Jun 22 • 8:36 AM EDT • Science • theconversation.com
From PCOS to PMOS: New name reflects women’s metabolic health

Global multidisciplinary health professionals and organizations have updated the name polycystic ovary syndrome (PCOS) to polyendocrine metabolic ovarian syndrome (PMOS), arguing that the former is “inaccurate.” They critique that the old name implied greater focus on pathological ovarian cysts, contributing to slower diagnosis, stigma, and limited research and policy framing for the 170 million women living with the condition.

nutritioninsight.com logo
Jun 15 • 4:14 AM EDT • Health • nutritioninsight.com
😐 Neutral (10%)
ahah
52%
wow
48%
USA Today columnist says United States has already 'lost' the World Cup because of politics, Donald Trump

USA Today columnist Nancy Armour claims the 2026 FIFA World Cup reveals America as a hateful nation, inaccurately blaming Trump and ticket prices.

foxnews.com logo
Jun 11 • 9:06 PM EDT • Politics • foxnews.com
Deep dive
😐 Neutral (9%)
wow
50%
ahah
25%
love
25%
Massachusetts attorney general’s lawsuit alleges $100M fraud by UnitedHealthcare

AG claims insurer pursued “growth-at-all-costs.” United calls the litigation “meritless” and inaccurate.

startribune.com logo
Jun 2 • 7:12 PM EDT • Health • startribune.com
Fact check: Colorado governor’s misleading rationale for freeing election denier Tina Peters

Colorado Gov. Jared Polis justified his decision to release election denier Tina Peters from prison with a series of false and misleading claims, inaccurately distancing her case from efforts to undermine the 2020 election.

cnn.com logo
May 18 • 7:38 PM EDT • Politics • cnn.com
The cinema effect: Turning films into a gateway to science

The sci-fi film Project Hail Mary, currently in theaters, is capturing the attention of both audiences and the scientific community for its science-based content. It manages to engage viewers with complex, cutting-edge topics — from astrophysics to language — without sacrificing entertainment. Yet not all films strike this balance. Many, unlike this recent production, have promoted inaccurate or even misleading scientific ideas, and thanks to their wide reach, have contributed to shaping distorted public perceptions of science. Precisely by harnessing the power of cinematic storytelling, Hildrun Walter, Fritz Treiber and colleagues have used films as a tool to foster engagement and dialogue between scientists and the public. Their innovative format, Science & Cinema, described in a new paper published in JCOM, proves both effective and potentially easy to replicate.

eurekalert.org logo
May 12 • 8:00 PM EDT • Science • eurekalert.org
In wild late-night posting spree, Trump attacks Obama with imaginary quote and false conspiracy theories

On Monday night, Pres. Donald Trump embarked on one of his periodic late-night social media posting sprees. As usual, his dozens of posts and re-posts were littered with debunked conspiracy theories and other wildly inaccurate claims

cnn.com logo
May 12 • 3:23 PM EDT • Politics • cnn.com
Deep dive
☹️ Negative (-3%)
ahah
64%
wow
36%
'The name was inaccurate': PCOS gets a new name after years-long effort

Polycystic ovary syndrome, or PCOS, has just been given a new name that experts say better reflects the nature of the condition.

livescience.com logo
May 12 • 5:00 AM EDT • Science • livescience.com
ESPN analyst misses mark in recent Jets roster dissection

The Jets may not have an elite receiving corps on paper, but calling wide receiver their biggest weakness feels inaccurate.

sports.yahoo.com logo
May 9 • 9:45 AM EDT • Sports • sports.yahoo.com
Top C.D.C. Official Delays Report on Covid Shot’s Effectiveness

Dr. Jay Bhattacharya objected to the study’s methodology, saying it gave an inaccurate picture of the vaccine’s benefits.

nytimes.com logo
Apr 9 • 10:54 AM EDT • Health • nytimes.com
Your stereotypes about Gen Z aren’t just inaccurate. They’re hurting your business

Welcome to Fast Company Daily, our daily newsletter on LinkedIn, featuring a free article selected each day by our editors as well as a roundup of great advice on careers, hiring, innovation, and technology. Visit fastcompany.

linkedin.com logo
Mar 27 • 10:30 AM EDT • Business • linkedin.com
Deep dive
😐 Neutral (15%)
wow
75%
ahah
25%
SoFi Responds to Inaccurate Short Seller Report

SoFi Technologies, Inc. (NASDAQ: SOFI), the one-stop shop for digital financial services, today issued the following statement: The claims made in the Muddy ...

businesswire.com logo
Mar 17 • 7:00 PM EDT • Business • businesswire.com
Trump T1 Phone: Specs and price proved inaccurate as manufacturer reveals final design

The Trump T1 Phone has been available for preorder since last June, but early buyers who’ve already paid a deposit may end up with a device that not only looks vastly different but also features altered specs compared with what was originally advertised. Trump Mobile has finally revealed the phone’s current design.

notebookcheck.net logo
Feb 9 • 8:01 AM EST • Technology • notebookcheck.net
A vaccine trial is called 'unethical' and a 'unique' opportunity. Is it on or off?

The U.S. is giving $1.6 million to researchers to study how the hepatitis B vaccine affects newborns in Guinea-Bissau. Local officials say the trial is suspended. U.S. officials say that's inaccurate.

npr.org logo
Jan 22 • 12:36 PM EST • Health • npr.org
High-Signal Publisher
Fact check: Trump’s barrage of false claims in Davos about Greenland and NATO

President Donald Trump’s Wednesday speech at the World Economic Forum in Switzerland was filled with inaccurate claims – notably including false and misleading statements about NATO and Greenland, the self-governing Danish territory he is pushing for the US to acquire.

cnn.com logo
Jan 21 • 11:44 AM EST • Politics • cnn.com
Deep dive
😐 Neutral (12%)
ahah
70%
love
30%
‘Flawed,’ ‘inaccurate,’ ‘biased’: Influential Oregonians say state’s lawyers wrote shoddy explanations for proposed ballot measure

The biggest flaw, they say, is that the draft explanations completely omit the proposed measures' biggest impact: 1.1 million more Oregonians would be able to take part in government funded primary elections for the most important offices.

oregonlive.com logo
Jan 13 • 5:39 PM EST • Politics • oregonlive.com
Deep dive
😐 Neutral (12%)
wow
51%
ahah
25%
love
24%
‘Dangerous and alarming’: Google removes some of its AI summaries after users’ health put at risk

Guardian investigation finds AI Overviews provided inaccurate and false information when queried over blood tests

theguardian.com logo
Jan 11 • 2:00 AM EST • Health • theguardian.com
Deep dive
😐 Neutral (5%)
ahah
55%
wow
45%
Google AI Overviews put people at risk of harm with misleading health advice

Exclusive: Inaccurate information presented in summaries, Guardian investigation finds

theguardian.com logo
Jan 2 • 12:00 PM EST • Health • theguardian.com
Inaccurate Fallout AI Recaps Pulled By Prime Video After Major Backlash

Backlash over inaccurate Fallout recaps causes backlash among viewers.

screenrant.com logo
Dec 12, 2025 • 10:44 PM EST • Technology • screenrant.com
AI chatbots used inaccurate information to change people’s political opinions, study finds

A paper published in the journal Science found that chatbots become more persuasive when they share large amounts of information.

nbcnews.com logo
Dec 4, 2025 • 2:00 PM EST • Politics • nbcnews.com
Deep dive High-Signal Publisher
😐 Neutral (15%)
wow
64%
ahah
36%
Chatbots can sway political opinions but are ‘substantially’ inaccurate, study finds

‘Information-dense’ AI responses are most persuasive but these tend to be less accurate, says security report

theguardian.com logo
Dec 4, 2025 • 2:00 PM EST • Politics • theguardian.com
Why YouTube Recap flopped and Spotify Wrapped is buzzing

YouTube launched its first Recap, but users say it's inaccurate and barebones. Spotify Wrapped, meanwhile, dominated with new features.

businessinsider.com logo
Dec 4, 2025 • 12:48 AM EST • Business • businessinsider.com
🙂 Positive (50%)