The implementation of evidence-based policies hinges on the dissemination of evidence to policymakers, a process influenced by the attributes of the sender. We conduct a country-wide RCT in which two ideologically opposite prominent think tanks, two major newspapers, and a research institution with nonsalient ideology communicate identical information about a low-cost, non-ideological, and effective policy based on published research findings to a large sample of Spanish local policymakers. We measure the impact of information directly on policy adoption and find heterogeneous effects. When the informing institution aligns ideologically with policymakers, communicating research results leads to a more than 65% increase in policy adoption compared to an uninformed control group, while informing from an opposite ideology does not lead to policy adoption. Our design also allows us to compare the impact of knowledge brokers, such as think tanks, and coverage in leading newspapers in adopting public policies. We find that, when ideologically aligned with policymakers, both are equally effective in increasing policy adoption. We propose a three-stage conceptual framework of policy adoption processes - selective exposure to information, belief updating, and policy implementation- and show that ideological alignment does not influence selective exposure to information. However, evidence from a post-intervention online experiment shows that ideological alignment affects belief updating regarding a recommended policy’s effectiveness. Finally, we discuss the trade-offs between effectiveness and outreach when using ideologically aligned and nonsalient institutions to disseminate research evidence and comment on the economic impact of ideological alignment for policy implementation.
Competitive selection processes can lead to inefficiencies in the labor market when disparities in performance during the selection process are unrelated to differences in job performance among selected candidates. Using data on the universe of candidates in the highly competitive and high stakes mandatory national entry exam into the medical profession in Spain over the past four decades, we first report the evolution of gender differences in exam performance, which translate into important gender gaps in the likelihood of securing a position (ranging from negative 7% up to positive 9% depending on the period), controlling for individual heterogeneity in ability. We then exploit the large variation in the proportion of available positions with respect to the number of candidates, to show that the observed evolution of gender gaps mimics the evolution of the selection process in terms of competitiveness: the more competitive the process, the higher the underperformance of women compared to men, while when the process shows low competitiveness, women outperform men. Aligning the number of candidates with the available positions based on the system’s needs rather than relying on alternative criteria would yield substantial gains in efficiency, particularly in professions where competitiveness does not play a pivotal role.
Online First: https://doi.org/10.1086/736551
Multiple choice question tests are often the gateway to important professional outcomes. We study gender differences in willingness to guess among highly skilled and trained candidates, who take a high stakes multiple choice question test, before and after a reduction in the number of alternative answers to each question, which sets the penalty for incorrect answers at the critical value. We find heterogeneous gender differences. We replicate the previous finding that women answer fewer questions than men and find that the reduction in the number of alternative answers levels the field for men and women but only among those candidates that answer most of the questions.
The burden of obesity, particularly severe and moderate forms, places significant strain on healthcare systems, mainly as it is concentrated amongst lower income individuals. In this study, we use food purchasing data of the whole population of low-income individuals who were beneficiaries in 2023 of vouchers provided by Red Cross Catalonia for food to conduct a randomised control trial in which participants receive either informational interventions about the benefits of healthy eating, “affective” nudges reminding facts about healthy food, or additional cash-aids. We find a positive impact of all interventions aggregated in two out of three standard indexes of food purchasing quality. Examining intervention heterogeneity, we find that informational interventions and affective nudges messages were more effective than cash in leading to healthier food choices.
DOI: https://doi.org/10.1016/j.socscimed.2025.118315
This paper serves as the opening for the Virtual Special Issue on the economic consequences of gender differences in behavior, published in the Journal of Economic Psychology. The issue aims to consolidate recent research exploring how gender differences in behavior, reflected in risk attitudes, competitiveness, and negotiation tendencies, impact economic outcomes and, in particular, may partially explain labor market differences across genders regarding occupational segregation, wage gaps, and disparities in career advancement. We provide an overview of the topic, highlighting the importance of understanding gender-specific economic behaviors and setting the stage for the detailed studies, and follow by introducing the articles included in the special issue, which use mainly experimental and empirical insights.
Online First DOI: https://doi.org/10.1016/j.joep.2025.102818
Standardized assessments are widely used to determine access to educational resources with important consequences for later economic outcomes in life and to evaluate and compare educational systems. However, many design features of the tests themselves may lead to psychological reactions influencing performance. In particular, the level of difficulty of the earlier questions in a test may affect performance in later questions. How should we order test questions according to their level of difficulty such that test performance offers an accurate assessment of the test taker's aptitudes and knowledge? We conduct a field experiment with about 19,000 participants in collaboration with an online teaching platform where we randomly assign participants to different orders of difficulty and we find that ordering the questions from easiest to most difficult yields the lowest probability to abandon the test, as well as the highest number of correct answers. We obtain similar results when exploiting the random variation of difficulty across test booklets in the Programme for International Student Assessment (PISA) for the years 2009, 2012, and 2015, which provides additional external validity to the experiment. We conclude that question difficulty order in tests has important policy implications for optimal test design and performance. It additionally may have important implications for ranking candidates, as well as for the evaluation and comparison of educational institutions and systems.
Multiple-choice tests are extensively used to measure individuals’ knowledge and aptitudes. We study gender differences in willingness to guess using approximately 10,000 multiple-choice math tests, where, for all participants, in half of the questions, omitted answers were rewarded while for the other half they scored the same as wrong answers. Using a within-participant regression analysis, we show that female participants leave significantly more omitted questions than males when there is a reward for omitted questions. This gender difference, which is stronger among high ability and older participants, hurts female performance as measured by the final score and position in the ranking. We conclude that it is important to use gender neutral scoring rules that do not differentiate between wrong answers and omitted questions in order to accurately measure individuals’ knowledge and aptitudes.
In two-stage elimination math contests participants from four different age groups compete to pass from stage 1 to stage 2 and later to be mong the winners. Although female participants have higher Math grades at school the gender gap reverses in the two stages of the contests. More importantly, following the same individual participant across different stages, we find that the gender gap in performance increases from stage 1 to stage 2 of the competition. The increase in female underperformance is attributed to higher competitive pressure and alternative explanations based on selection, discrimination and differences in reaction to increasing difficulty are ruled out.
We show that the existence of gender differences in performance is highly sensitive to the task used to measure performance, to existing stereotypes and to informational conditions. Out of sixteen purposely designed treatments we find that women underperform when competing only when two conditions are met: 1) the task used is perceived as favoring men and 2) the presence of a rival is strongly primed through information provided before competing. Such sensitivity sheds light on the contradictory evidence found on stereotype-threat causing gender differences in performance under competition.
We report evidence from a large field experiment that compares the effectiveness of contingent and noncontingent incentives in eliciting costly effort for a large range of payment levels. The company with which we worked sent 7,250 letters asking customers to complete a survey. Some letters promised to pay amounts ranging from $1 to $30 upon compliance (contingent incentives), whereas others already contained the money in the request envelopes (noncontingent incentives). Compared to no payment, very small contingent payments lower the response rate while small noncontingent payments raise the response rate. As expected, response rates rise with the size of the incentive offered. The response rate in the noncontingent incentives rises more rapidly for low amounts of incentive, but then flattens out and reaches lower levels than under contingent payments. We discuss how the optimal policy regarding the use of each size and type of incentives crucially depends on firms’ objectives.
Using data from modified dictator games and a mixture-of-types estimation technique, we find a clear relationship between a classification of subjects into four different types of interdependent preferences (selfish, social welfare maximizers, inequity averse, and competitive) and the beliefs subjects hold about others’ distributive choices in a nonstrategic environment. In particular, selfish individuals fall into false-consensus bias more than other types, as they can hardly conceive that other individuals incur costs so as to change the distribution of payoffs. We also find that selfish individuals are the most robust preference type when repeating play, both when they learn about others’ previous choices (social information) and when they do not, while other preference types are more unstable.
Affirmative action policies bias tournament rules in order to provide equal opportunities to a group of competitors who have a disadvantage they cannot be held responsible for. Its implementation affects the underlying incentive structure which might induce lower performance by participants, and additionally result in a selected pool of tournament winners that is less efficient. In this paper, we study the empirical validity of such concerns in a case where the disadvantage affects capacities to compete. We conducted real-effort tournaments between pairs of children from two similar schools who systematically differed in how much training they received ex-ante on the task at hand. Contrary to the expressed concerns, our results show that the implementation of affirmative action did not result in a significant performance loss for either advantaged or disadvantaged subjects; instead it rather enhanced the performance for a large group of participants. Moreover, affirmative action resulted in a more equitable tournament winner pool where half of the selected tournament winners came from the originally disadvantaged group. Hence, the negative selection effects due to the biased tournament rules were (at least partially) offset by performance enhancing incentive effects.
We present evidence from an experimentin which groups select a leader to compete against the leaders of other groups in a real-effort task that they have all performed in the past. We find that women are selected much less often as leaders than is suggested by their individual past performance. We study three potential explanations for the underrepresentation of women, namely, gender differences in overconfidence concerning past performance, in the willingness to exaggerate past performance to the group, and in the reaction to monetary incentives. We find that men’s overconfidence is the driving force behind the observed prevalence of male representation.
We compare behavior in modified dictator games with and without role uncertainty. Subjects choose between a selfish action, a costly surplus creating action (altruistic behavior) and a costly surplus destroying action (spiteful behavior). While costly surplus creating actions are the most frequent under role uncertainty (64%), selfish actions become the most frequent without role uncertainty (69%). Also, the frequency of surplus destroying choices is negligible with role uncertainty (1%) but not so without it (11%). A classification of subjects into four different types of interdependent preferences (Selfish, Social Welfare maximizing, Inequity Averse and Competitive) shows that the use of role uncertainty overestimates the prevalence of Social Welfare maximizing preferences in the subject population (from 74% with role uncertainty to 21% without it) and underestimates Selfish and Inequity Averse preferences. An additional treatment, in which subjects undertake an understanding test before participating in the experiment with role uncertainty, shows that the vast majority of subjects (93%) correctly understand the payoff mechanism with role uncertainty, but yet surplus creating actions were most frequent. Our results warn against the use of role uncertainty in experiments that aim to measure the prevalence of interdependent preferences.
This paper uses subjects’ diverse self-reported justifications to explain discrepancies between observed heterogeneous behavior and the unique equilibrium prediction in a oneshot traveler’s dilemma experiment. Principal components analysis suggests that iterative reasoning, aspiration levels, competitive behavior, attitudes towards risk and penalties and focal points may be behind different choices. Such reasons are coherent with same subjects’ behavior in other tests and experiments in which these particular issues are prominent, and thus, we identify “types” of subjects. Overall, we conclude that subjects’ self-justifications in complex strategic situations contain informational value which may be used to predict behavior in other situations of economic importance.
We ask whether the absence of information about other voters’ preferences allows optimal voting to be interpreted as sincere.We start by classifying voting mechanisms as simple and complex according to the number of message types voters can use to elect alternatives. We show that while in simple voting mechanisms the elimination of information about other voters’ preferences allows optimal voting to be interpreted as sincere, this is no longer always true for complex ones. In complex voting mechanisms, voters’ optimal strategy may vary with the size of the electorate. Therefore, in order to interpret optimal voting as sincere for complex voting mechanisms, we describe the optimal voting strategy when voters not only have no information but also have no pivotal power, i.e., as the size of the electorate tends to infinity.
We report experimental results on a series of ten one-shot two-person 3 × 3 normal form games with unique equilibrium in pure strategies played by non-economists. In contrast to previous experiments in which game theory predictions fail dramatically, a majority of actions taken coincided with the equilibrium prediction (70.2%) and were best-responses to subjects’ stated beliefs (67.2%). In constant-sum games, 78% of actions taken were predicted by the equilibrium model, outperforming simple K-level reasoning
models. We discuss how non-trivial game characteristics related to risk aversion, efficiency concerns and social preferences may affect the predictive value of different models in simple normal form games.
We study optimal contracts in a simple model where employees are averse to inequity, as modeled by Fehr and Schmidt (1999). A “selfish” employer can profitably exploit envy or
guilt by offering contracts which create inequity off-equilibrium, i.e., when employees do not meet his demands. Such contracts resemble team and relative performance contracts. We derive conditions for inequity aversion to be in itself a reason to form work teams of distributionally concerned employees, even in situations in which effort is contractible.
In this paper we study the mechanics of “leading by example” in teams. Leadership is beneficial for the entire team when agents are conformists, i.e., dislike effort differentials. We also show how leadership can arise endogenously and discuss what type of leader benefits a team most.
At the beginning of the 1990s it appeared that there was considerable agreement about the kind of economic policies that countries turning to the IMF and the World Bank should pursue. These included macroeconomic stabilisation, microeconomic liberalisation and openness, and were summarised by the concept of a ‘Washington Consensus’. How has the Consensus stood up to the passage of time? This article briefly assesses the track record of Consensus-type policies and shows how the Consensus has evolved. With regards to some of its components, a greater sense of agnosticism may now prevail. Moreover, issues that were little or no part of the Consensus have come to the fore. The implications of these changes for institutional design are also investigated
Identifying higher orders of rationality is crucial to the understanding of strategic behavior. Nonetheless, the identification of a subject’s actual order of rationality from observed behavior in games remains highly elusive. Games may significantly impact and hence invalidate the identified order. To tackle this fundamental problem, we introduce an axiomatic approach that singles out a new class of games, the e-ring games. We then present results from a within subject experiment comparing individuals’ classification across e-ring games and standard games previously used in the literature. The results show that satisfying the axioms introduced significantly reduces errors and contributes towards a more reliable identification.
Active labor market programs (ALMPs) are widely used to boost employment prospects for individuals at risk of social exclusion, though evidence on their effectiveness remains mixed. While most studies focus on single interventions, less is known about the impact of combining multiple ALMPs that may complement each other to enhance both employability and psychological well-being. This paper evaluates the effectiveness of a comprehensive ALMP bundle delivered in partnership with Red Cross Spain to facilitate reinsertion in the labor market. Participants were randomly assigned to a control group with access to Red Cross’s regular services or to a treatment group that received additional interventions, including psychological support, soft skills development, digital literacy, certified professional courses, and job search assistance. Using administrative data, we find no significant effects of the treatment on employment outcomes immediately after the intervention, or six and twelve months later. However, the program does improve certain outcomes related to employability and personal autonomy. Specifically, it reduces the risk of poor psychological health and enhances frustration tolerance, problem identification,
knowledge of available resources, and digital skills. Participants also report a higher perceived likelihood of finding a job and use a broader range of job search methods.
Understanding what affects consumption satisfaction is fundamental to understanding consumer behavior. Measuring satisfaction, however, is not trivial, especially in the context of experience goods where perceived quality is often subjective and unobservable prior to consumption. We report the results of a field study (N = 433) conducted in collaboration with a theater that uses pay what-you-want (PWYW) pricing, inviting audience members to pay at the end of the show. Our analyses indicate that neither expected nor realized enjoyment predict payments once we control for the expectation-realization gap. The paper highlights the managerial implications of these findings.
Over the last decades, work tasks have become increasingly non-routine, complex, and analytical, leading to the widespread adoption of team-based organizational structures. To assemble productive teams and implement efficient governance structures, human resource (HR) experts need to form correct expectations about the most crucial determinants of team success. This study documents HR experts’ perceptions (n=3,000) regarding the relative importance of various team composition dimensions and governance structures for performance in non-routine analytical tasks. Exploiting the unique opportunity to contrast expectations with actual performance data of 1,062 teams, we show that experts hold qualitatively accurate beliefs. However, they substantially underestimate the value of leadership. These patterns hold up in an additional general population sample (n=3,000). Furthermore, we document implicit biases against (particularly female) leadership, which partially depend on the respondent’s own gender.
This pilot project implemented by Save the Children Spain, a non-governmental organization dedicated to promoting public policies that improve the lives of children and adolescents, seeks to evaluate the effect of providing a comprehensive program that combines social, educational and labor market integration interventions to promote the well-being of households that live socially excluded or at-risk of social exclusion. The intervention was carried out between September 2022 and September 2023 in four Spanish municipalities. The intervention design was based on the comparison of four treatments that were comprised of a stratified random allocation of households. The study finds impacts of the comprehensive treatment in a reduction of material and social deprivation, an increase in monthly household income, an improvement in parents' expectations of their children's studies, and an improvement in the standardized test scores for language and mathematics. Despite this, it is difficult to conclude that the comprehensive model proposed by STC is more effective in improving the well-being of households with children and adolescents who live socially excluded or at-risk of exclusion than traditional programs where only social support is provided. This is because no statistically significant impacts are observed in most of the main indicators for quality of life, employment and educational continuity of the participating households.
We study the unique role of social distance in determining deceptive behavior, independent of other confounding factors such as information asymmetries. In this experiment, confederates of three different ethnicities, under the same informational conditions, take bicycle taxi rides in Malawi, where drivers have two opportunities not to adhere to social norms of honesty - by overcharging or outright stealing by not returning back with change. We find that social distance affects fraudulent behavior per se, increasing the charged price by 11%. Using the outright stealing measure, we find that deceptive behavior is independent of stake size. Finally, we find evidence that moral priming affects cheating, as routes that have destinations with moral connotations, such as churches, hospitals, and schools present fraud 18% less frequently than routes without moral connotations and, in these routes, discrimination in overcharges among different ethnicities disappears. Our results suggest that in aid programs in developing countries, fraud may be reduced if the identity of the agents interacting with the locals is carefully chosen and the purposes of the aid are given proper moral connotations.
Paper available soon
We perform a further experiment to check the robustness of the main result in Rey Biel (2005) to sequential play. We find that Equilibrium predictions work even better when the same games are played sequentially: 85% of first movers choose the Equilibrium strategy and 85% of second movers best respond to the action taken by first movers. We conclude by identifying constant sum games as a class of games where experimental subjects' choices coincide with theory predictions and we argue that in such games distributional and reciprocal preferences do not influence subjects' decisions.
This paper tries to find an economic explanation to the fact that vaccines for death-causing diseases such as AIDS, malaria and tuberculosis, have not been discovered yet. We argue that private laboratories do not have proper incentives to invest in research due to three reasons: fi) the vast majority of demand comes from poor countries who cannot afford a price high enough to compensate research costs, £) international pressure to reduce vaccine prices once discovered will be high and successful, and 3) laboratories act as monopolists in the competing market of treatments for already infected patients and the appearance of a vaccine will, in the short run, lower the price of treatments while, in the longer run, make the treatment market disappear. We first present a static model to show why laboratories may under-invest in vaccines. We extend the model to a dynamic context to endogenize how the infection rate. Finally, we discuss some mechanisms to provide incentives for private research in vaccines.
We study how giving depends on income and luck, and how culture and information about the determinants of others’ income affect this relationship. Our data come from an experiment conducted in two countries, the US and Spain – each of which have different beliefs about how income inequality arises. We find that when individuals are informed about the determinants of income, there are no crosscultural differences in giving. When uninformed, however, Americans give less than the Spanish. This difference persists even after controlling for beliefs, personal characteristics, and values.
Economic theories and experiments could and should inform each other. An economic theory is more useful if it is not only an intellectual exercise but also relates to empirical relevant behavior. Experiments that are based on a set of alternative, well defined, hypotheses are more useful. We argue that these theories do not have to be a mathematical model. For example, experiments can help in the understanding and testing of mechanisms that are used in the world, but for which deriving equilibrium behavior analytically is too complicated.
The shift toward evidence-based policymaking also entails embracing research from the behavioral sciences, which has significantly contributed to this approach. After clarifying that behavioral interventions go beyond the nudge framework, we discuss the advantages and challenges of collaboration between public administrations and the research community in implementing policies that account for how citizens interpret them. We focus on four key requirements for successful interventions: a favorable attitude toward experimentation, public availability of administrative data, evaluable intervention design, and independent, rigorous policy evaluation.
Background. Immunization against preventable diseases as meningitis is crucial from a public health perspective to face challenges posed by these infections. Nurses hold a great responsibility for these programs, which highlights the importance of understanding their preferences and needs to improve the success of campaigns. This study aimed to investigate nurses' preferences regarding MenACWY conjugate vaccines commercialized in Spain.
Methods. A national-level discrete choice experiment (DCE) was conducted. A literature review and a focus group informed the DCE design. Six attributes were included: pharmaceutical form, coadministration evidence, shelf-life, package contents, single-doses per package and package volume. Conditional logit models (CLM) quantified preferences and relative importance (RI).
Findings. Thirty experienced primary care nurses participated in this study. Evidence of coadministration with other vaccines was the most important attribute (RI=43·78%), followed by package size (RI=22·17%), pharmaceutical form (RI=19·07%) and package content (RI=11·80%). There was a preference for evidence of coadministration with routine vaccines (odds ratio [OR]=2·579, 95% confidence interval [95%CI]=2·210-3·002), smaller volumes (OR=1·494, 95%CI=1·264-1·767), liquid formulations (OR=1·283, 95%CI=1·108-1·486) and package contents including only vial/s (OR=1·283, 95%CI=1·108-1·486). No statistical evidence was found for the remaining attributes.
Interpretation. Evidence of co-administration with routine vaccines, easy-to-store packages and fully liquid formulations were drivers of nurses’ preferences regarding MenACWY conjugate vaccines. These findings provide valuable insights for decisionmakers to optimize current campaigns.