1Policy success is now widely understood as a multidimensional concept, moving beyond narrow outcome-based assessments. Early approaches tended to equate success purely with the achievement of policy goals or the production of net positive effects, often framed in terms of effectiveness or cost–benefit considerations (Nagel, 1980). Over time, the literature has increasingly recognized that policy success cannot be reduced to programmatic performance alone. A growing body of work has expanded the concept to include additional dimensions, such as process, political, and, more recently, temporal or endurance-based success (Bovens et al., 2002; Marsh & McConnell, 2010; Newman, 2014; Compton & ’t Hart, 2019; Lindquist et al., 2022). These contributions highlight that policies must also be evaluated in terms of how they are formulated and adopted (process), how they are perceived and supported in the political arena (political), and whether their effects endure over time (endurance). As a result, policy success involves trade-offs across competing criteria, actors, and time horizons.
2While the concept of policy success has become increasingly complex through the addition of new dimensions, the programmatic dimension itself has remained comparatively under-theorized. Across the literature, it is generally approached in instrumental terms, as the extent to which policies achieve their intended objectives and produce desirable outcomes. Programmatic success is thus defined in terms of effectiveness and efficiency (Bovens et al., 2002), whether benefits exceed costs (Shuck, 2014), and further elaborated through criteria such as operational success in implementation, outcome achievement, and efficiency (Marsh & McConnell, 2010). This perspective is further reflected in McConnell’s (2010) broader definition of policy success, where a policy is considered successful if it achieves the goals set by its proponents and attracts little or no significant criticism. Even when distributional effects and goal ambiguity are acknowledged, authors retain the centrality of an outcome-based orientation founded in the stated policy objectives (Newman, 2014; Lindquist et al., 2022).
3As a result, complexity is primarily located across dimensions — between programmatic, process, and political success — rather than within the programmatic dimension itself. Programmatic success is therefore typically treated as the least problematic component, with outcomes implicitly assumed as given or derived from policy goals and problem definitions. The analytical challenge is thus framed as assessing whether these outcomes are achieved, rather than questioning how they are defined in the first place.
4Yet, the literature has long acknowledged that policy outcomes are not straightforward. Work on policy fiascoes has emphasized the substantial interpretive and conceptual challenges regarding what outcomes to consider, how to measure them, and when to observe them (Bovens & ’t Hart, 1998). Similarly, McConnell (2010) underscores that programmatic success is inherently contested, varying across actors, time horizons, and evaluative criteria, and involving trade-offs among multiple and sometimes conflicting outcomes.
5Despite these insights, such complexity has rarely been addressed as a prior problem of what counts as an outcome and how to conceptualize success. Where acknowledged, it is often framed in terms of multiple or contested interpretations and addressed at the stage of evaluation — after outcomes have already been defined — rather than as a prior problem of outcome specification. In other words, while the literature recognizes that outcomes may be ambiguous, contested, or politically constructed, it has paid comparatively limited attention to how they are identified and made observable in the first place.
6This article addresses this gap by returning to the programmatic dimension of policy success and arguing for the need to explicitly theorize outcomes in its appraisal. Specifically, before assessing whether a policy works or explaining how and why it produces effects, it is necessary to define what counts as an outcome in the first place. Programmatic success is not simply observed through given outcomes, but constructed through prior analytical choices that shape how it is subsequently measured, assessed, and interpreted.
7To capture this problem, the article identifies four key dimensions of outcome theorization: the outcome domain (what counts as an outcome), construct selection (which dimensions of change represent program effects), intermediate outcomes (where along the causal chain success is located), and timing (when outcomes become observable). Together, these dimensions identify the core analytical challenges involved in making programmatic success observable.
8To support outcome theorization, the article draws on theory-driven evaluation (TDE). By linking program activities to the mechanisms through which change is generated, TDE provides a basis for identifying relevant outcomes and understanding their underlying dynamics. The article therefore uses a theory-driven approach to specify and refine outcome sets prior to measurement, clarifying what counts as success and how it can be observed. As such, the approach can be understood as an upstream analytical step in program appraisal, supporting policy analysts and evaluators before appraising programmatic success.
9Rather than providing a direct assessment of program performance or substituting for ex post evaluation, the approach operates at the level of outcome specification, constraining and structuring what can be meaningfully evaluated and reducing the risk of mis-specification that may bias subsequent appraisal. Making these assumptions explicit enables a more systematic identification of outcomes, their appropriate conceptualization, and their positioning along causal and temporal dimensions.
- 1 A related, multi-author conference paper (Busetti et Al. 2025) develops and empirically tests an FB (...)
10The article proceeds as follows. The next section analyzes the four dimensions in the appraisal of programmatic success. The third section examines how theory-driven evaluation can help address these challenges and support outcome theorization. The methodology outlines the research design, data sources, and limitations of the empirical study. The empirical section illustrates this procedure through the case of food bank markets (FBMs), a model of food assistance that provides access to food through market-like settings. Despite the existence of widely accepted definitions of food security, the case shows how the identification of relevant outcomes remains analytically uncertain at the program level, requiring explicit theory-building to define what counts as success in specific contexts. The discussion section reflects on the main findings of the empirical analysis, while the conclusion outlines the broader theoretical implications of the study.1
11Assessing policy outcomes presents several hurdles. Effects may be unexpected, difficult to capture, or only surface after a long and winding sequence of events. This section identifies four challenges that influence how program success is defined and made observable: specifying the outcome domain beyond official goals (what counts as an outcome), selecting analytically suitable constructs within that domain (which effects can validly be used to appraise success), identifying intermediate outcomes that link program activities to their effects (where relevant changes occur along the causal chain), and determining when such outcomes should be expected to emerge, peak, or fade (when success becomes observable). Table 1 summarizes the four analytical challenges discussed in this section.
Table 1. Dimensions in the appraisal of programmatic success.
|
DIMENSION
|
Core question
|
Analytical focus
|
Risk if neglected
|
Illustrative exampleS (fbm)
|
|
Outcome domain
|
What counts as an outcome?
|
Defining the set of effects plausibly associated with the program
|
Relevant effects may be excluded or irrelevant ones included
|
Distinguishing food waste as a by-product of the donation systems rather than a program outcome
|
|
Construct selection
|
Which outcomes can validly represent program success?
|
Selecting constructs capturing program-relevant dimensions of change
|
Program success may be misrepresented if constructs are poorly aligned with program mechanisms
|
Interpreting food security with caution, given its limited sensitivity to FBM interventions
|
|
Intermediate outcomes
|
Where along the causal chain should success be located?
|
Identifying whether and which intermediate outcomes are relevant
|
Success may be underestimated (if based on distal outcomes only) or core processes may be ignored
|
Including food procurement capacity and users’ behavior as relevant dimensions of program appraisal
|
|
Timing
|
When should outcomes be observed?
|
Accounting for how outcomes emerge and vary over time
|
Effects may be missed, overstated, or misinterpreted
|
Recognizing that social inclusion may develop gradually and take a longer time to be observable
|
Source: the Author
12Public policy typically includes official goals that guide the appraisal of programmatic success and the search for outcomes. However, official goals are often incomplete representations of program effects; they are formulated in generic terms (Weiss, 1972) and are typically selected based on desirability rather than possibility (Chen & Rossi, 1980). Policy goals are vague, set unrealistically high or low, continually adjusted, and may be at odds with one another (Bovens & ’t Hart, 1998). Put simply, they can be better understood at the political level rather than as measures of a program’s performance.
13Accordingly, understanding program success requires explicit work to define a plausible outcome domain, namely, the set of effects that may arise in connection with the program. Investigating how the program is supposed to work may shed light on which official goals are pertinent to the program as implemented, rather than being symbolic or too generic, and whether additional outcomes or side effects may arise from how the intervention operates.
14For example, Marchildon (2022) characterizes Canadian Medicare’s official objective as ensuring access to health services by removing financial barriers. However, programmatic success was also assessed across other dimensions, such as the cost advantages for Canadian businesses in not having to insure workers for the most expensive treatments, increased administrative efficiency, and the feedback effect encouraging more primary and preventative care, ultimately reducing downstream treatments (Marchildon, 2022). As another example, a sanction-based program implemented in Italy had the official goal of reducing drug consumption through threat and punitive measures (Leone, 2008). Yet, some sanctioned individuals reacted by acquiring strategic behaviors, i.e., not reducing drug use but becoming more careful about where and when to consume drugs. In other cases, the program produced positive spillover effects: friends of the sanctioned individuals became more reflective and were deterred from consumption.
15These examples illustrate how program effects extend beyond officially stated goals, encompassing indirect or unintended consequences. Such effects expand the outcome domain relevant to appraising program success and can be identified through a combination of theoretical reasoning about the program’s logic and exploratory empirical research. While theory cannot anticipate all unintended consequences, it can provide guidance on how actors’ behaviors, capacities, and program contexts may generate spillovers or secondary outcomes.
16Issues of construct validity and measurement have long been addressed in both quantitative and qualitative methodologies (Cronbach & Meehl, 1955; Goertz, 2006), primarily concerning the attributes, operationalization, and validation of constructs once they have already been identified as relevant. Our challenge focuses instead on an earlier stage: the theorization and specification of which outcome constructs should be considered first when appraising program success. This step involves moving from an inclusive outcome domain of plausible intended and unintended effects to a selective set of constructs best suited to the valid appraisal of program success.
17Programs operate through specific activities and mechanisms that affect only particular dimensions of a policy problem. Even when multiple plausible outcomes have been identified within the outcome domain, outcome constructs must be selected in relation to this theoretical causal logic: what the program actually does and how change is expected to occur (Bickman, 1987). If this step is neglected, constructs are selected for convenience, data availability, or by generic analogy with similar programs, and may misrepresent program success or fail to capture the specific changes generated by the program.
18For example, in their evaluation of health care services, Greenhalgh et al. (2009) describe how the introduction of standard guidelines among health care providers was framed as a means of providing consistent service and care for patients. However, their research found that some health-care providers had already developed local guidelines tailored to local needs and the demands of local minority populations. While the program could directly influence the adoption of uniform standards (making standardization part of the outcome domain), using standardization as the primary construct for assessing consistency of care risked misrepresenting program success, as it failed to capture the program’s contribution to context-sensitive adaptations.
19In other cases, different constructs capture distinct dimensions of the same policy goals, depending on how they relate to program activities. For instance, in the case of food security, alternative constructs include food access, dietary variety, or the ability to tailor food choices to individual preferences. However, these dimensions do not equally align with all programs. An intervention providing standardized food parcels may improve access, while having limited effects on dietary autonomy. Selecting among these constructs within food security is analytically prior to defining how each will be measured (e.g., economic versus physical access) and directly shapes the interpretation of program success.
20Finally, the choice of outcome constructs also affects the ability to detect changes. Lipsey (1990), for example, compared alternative constructs of language performance among children enrolled in special education classes and found that they yielded different effect sizes. IQ tests, for example, showed no effect, precisely because IQ is a relatively stable trait that program activities were unlikely to influence. In this case, the construct itself did not align with the program’s causal logic and was poorly suited to detecting change.
21A further dimension concerns where along the causal chain program success should be located, and whether intermediate outcomes should be treated as relevant dimensions of program performance. Policy interventions are ultimately directed at distal outcomes such as reducing poverty, preventing crime, or improving healthcare. In some cases, these outcomes follow directly from program outputs, the immediate results of program activities. Often, however, program effects unfold through longer chains of intermediate outcomes, i.e. changes in the behavior, perceptions, or capacities of relevant actors triggered by program outputs and activities.
22While distal outcomes are substantively and normatively meaningful, they are often weakly attributable and exposed to confounding influences. In contrast, intermediate outcomes are more directly linked to program activities and more sensitive to its effects, although they capture only partial dimensions of success. Crucially, however, intermediate outcomes are not merely early indicators of distal changes, but may constitute core dimensions of how the program works.
23In this respect, theory-driven evaluation (Chen, 2015; Donaldson, 2022) highlights the importance of identifying intermediate outcomes as part of analyzing how programs achieve their expected social benefits and where implementation problems may arise (Lipsey & Pollard, 1989). From the perspective of outcome appraisal, this implies that program success should not be assessed solely in terms of final outcomes, but may also be meaningfully evaluated through the intermediate processes that sustain or constrain program effects.
24Consider, for instance, a mentoring program aimed at improving school performance among at-risk youth. Depending on the characteristics of the program, improved grades may be a valid distal outcome measure of program success. Still, this result may only emerge after a series of intermediate steps: improved mentee’s perception of scholastic competence, increased value attributed to school, higher quality in parental relationship, and ultimately school performance (Rhodes et al., 2000). Appraising program success only in terms of academic performance may obscure not only early progress, but also the core processes through which the program operates.
25A further dimension concerns when outcomes (proximal or distal) should be expected to appear and how their temporal dynamics affect their observability. The timing of outcomes is not necessarily linear or stable. Outcomes may develop gradually, appear only after a threshold of exposure, immediately peak but quickly decay, or fluctuate as different mechanisms become active over time. This introduces ambiguity in appraising programmatic success: measuring too early may underestimate program effectiveness when changes have not yet materialized, while measuring too late may obscure effects that are temporary, have dissipated, or have been overtaken by other intervening processes. The same outcome may therefore yield different — and potentially conflicting — assessments of program success depending on when it is observed. Without explicit attention to these temporal dynamics, such variation risks being misinterpreted as inconsistencies in results or a lack of effect.
26To take one example, in his analysis of street lighting interventions, Pawson (2006) shows how crime levels fluctuated sharply across different phases of the program — rising, falling, and rising again — without a stable pattern. These shifts reflect the layered interaction of multiple mechanisms, each becoming active at different times: increased visibility, greater foot traffic, community confidence, and more. Without a theory to guide expectations about when these mechanisms might work, assessing program success at a single moment may provide a partial or misleading picture of its effects.
27The four analytical challenges outlined above highlight a common underlying issue: outcomes cannot be separated from assumptions about how programs generate effects. Addressing these challenges, therefore, requires an analytical basis for linking program activities to the processes through which change occurs, and to the outcomes through which success is assessed.
28Theory-driven evaluation (TDE) provides such a foundation. TDE refers to a family of approaches in evaluation practice that seek to explain program effects by making explicit the mechanisms linking program activities to outcomes (Chen & Rossi, 1980, 1983; Weiss, 1997; Funnell & Rogers, 2011). Rather than focusing solely on whether a program works, these approaches investigate how, why, and for whom change occurs (Pawson & Tilley, 1997).
29Early developments in TDE already highlighted that outcomes cannot be treated as ready-made objects of evaluation. In 1972, Peter H. Rossi formulated the ‘iron law’ of evaluation (Rossi, 1987, 2003), highlighting that the persistent absence of measurable effects in impact assessments was often linked to the weak theoretical grounding of the programs under study rather than to program failure - a distinction that reframes the absence of effects as a problem of outcome specification before it is a problem of program performance. Similarly, ‘goal-free evaluation’ highlighted the limits of restricting analysis to intended outcomes, emphasizing the opportunity to increase program learning by considering a broader range of effects arising from program operations (Scriven, 1972). Building on these concerns, Chen and Rossi initially proposed TDE as a ‘multi-goal’ approach in which, prior to measurement, evaluators should define a broad set of potential outcomes informed by theory and empirical knowledge (Chen & Rossi, 1980).
30Although subsequent developments in TDE have primarily focused on explaining why and how outcomes emerge rather than on systematically examining their theoretical grounding (Coryn et al., 2011; Salter & Kothari, 2014), the perspective provides a useful basis for outcome theorization. The reason is that program causality and outcome specification are not independent analytical steps - they are constitutively linked. As Glennan (2002) argues, mechanisms are always mechanisms for specific behaviors: how a system is understood to work depends on which behavior is under consideration, and vice versa. In program terms, this means that asking how a program produces change and asking what counts as a relevant outcome are not sequential questions. Specifying what program activities do, and through what processes, simultaneously delimits which effects are plausibly within the program’s reach and which constructs can meaningfully capture them.
31This perspective maps directly onto the four dimensions identified above. Mechanisms define what belongs in the outcome domain and guide construct selection toward effects that are sensitive to how change actually occurs; their activation along the causal chain makes intermediate outcomes analytically central; and their temporal dynamics — which may unfold gradually, non-linearly, or in interaction with context — determine when effects become observable.
32These insights, however, have not significantly informed the policy literature on programmatic success, which tends to engage with evaluation primarily in its impact- and performance-oriented approaches. Within these strands, outcomes are typically treated as predefined objects of verification and are often equated to goal attainment - a tendency noted by Marsh and McConnell (2010, p. 565).
33Following TDE’s original multi-goal proposals, this article conceives outcome theorization as a preliminary analytical step informed by mechanism logic, prior to developing a full program theory or testing causal relationships. This means drawing on TDE reasoning to ask what effects program activities can plausibly generate, not by formally mapping mechanisms, but by using the logic of how programs work to discipline the identification and selection of outcomes. The approach does not require a complete causal model, but rather sufficient theoretical grounding to distinguish effects within the program’s causal reach from those that are not, and to identify the constructs through which success can be meaningfully assessed.
34Treating outcome specification as a preliminary analytical step is particularly relevant to evaluation design and early-stage program appraisal, where policy analysts and evaluators need to define what constitutes a meaningful outcome before selecting measurement strategies. In this perspective, the outcome domain, construct selection, intermediate outcomes, and timing can be understood as the four analytical dimensions through which this specification is operationalized, as developed in the methodology section that follows.
35Case selection. The empirical analysis focuses on three food bank markets (FBMs) managed by local voluntary associations within the Italian Red Cross. This selection ensures comparability in organizational procedures, volunteer training, and national guidelines, while allowing variation in food provision, access procedures, and relationships with municipalities and donor networks. The three FBMs were selected in consultation with the Italian Red Cross based on three criteria: operational maturity (more than three years of activity), prior experience with food parcel distribution (allowing comparison across models), and recognized operational capacity based on the national Red Cross’s knowledge of their ongoing activity. The case selection is not aimed at representativeness, but at identifying analytically relevant cases while introducing variation in key program features in order to maximize the range of observable outcomes.
36Food aid provides a suitable empirical context for analyzing program outcomes. Food security is commonly defined as stable access to sufficient, safe, and nutritious food that meets people’s dietary needs and preferences (FAO, 2001). This widely accepted definition offers a clear reference point for examining which outcomes are included, how they are specified, and which dimensions are selected in practice.
37Data sources. The study draws on two sources: (1) secondary evidence, including program documents and academic and grey literature, and (2) six exploratory interviews with key informants (three volunteers, one policymaker, and two representatives of the Italian Red Cross). Interviews were conducted in person and online and lasted approximately one hour. Interviews began with open-ended questions aimed at eliciting expected and observed program effects, followed by targeted follow-up questions to assess and refine the initial hypotheses.
38Limitations. The analysis focuses on outcome specification rather than outcome testing. As such, the findings do not aim to provide evidence on program effectiveness, but to identify and theoretically ground outcome constructs that will drive the appraisal and measurement of program success. The limited number of interviews and restricted geographic scope constrain the ability to capture the full range of contextual variation across FBMs; therefore, further empirical work is required to assess the validity and generalizability of the specified outcomes across settings. These limitations are inherent to outcome specification, which necessarily constitutes a preliminary stage in the appraisal of program success.
39Research design and outcome specification. The study uses a realist review, a technique within the theory-driven approach for collecting and synthesizing secondary evidence with an explanatory focus on why and how a program works (Pawson, 2002; Pawson et al., 2004), complemented by exploratory interviews. Outcomes are defined as changes in users’ conditions, behaviors, or capacities that can plausibly be linked to program activities. This definition distinguishes outcomes from outputs, understood as the direct deliverables of program activities (e.g. food provision), and focuses instead on changes emerging from the interaction with program design features.
40The specification of outcomes followed an iterative, theory-informed process articulated in three analytical phases: initial identification and extraction, empirical qualification and refinement, and outcome specification and consolidation.
-
Identification and initial extraction. Outcome identification starts from program documents and official goals, combined with theory-driven reasoning on program features. Drawing on TDE logic, program activities are analyzed in terms of their causal powers and the mechanisms through which they may plausibly generate change. By linking program activities and design features to potential effects, this phase may signal causally relevant design and implementation features. This phase yields an initial set of candidate outcomes and deductive hypotheses regarding supporting program activities and intermediate outcomes.
-
Empirical qualification and refinement. The initial outcome set is then qualified and expanded through document and literature review and exploratory interviews. This phase is used to assess the empirical plausibility of candidate outcomes and their most suitable constructs, identify supporting or constraining program and contextual conditions, and refine the outcome set by distinguishing between final and intermediate outcomes. It also enables the identification of temporal dynamics that affect outcome observability.
-
Outcome specification and consolidation. Finally, evidence from previous steps is synthesized to specify the outcome set. This involves selecting, redefining, or excluding candidate outcomes, clarifying their relationship to program activities, and determining whether they function as final or intermediate outcomes. The result is a refined theory-informed outcome specification. In the case of the FBM, for instance, the process led to the requalification of three outcomes (food security, broader well-being, and social inclusion), the inclusion of intermediate outcomes, and the exclusion of one outcome (food waste).
41Figure 1 illustrates this process. It shows how each outcome was empirically derived step-by-step: starting from official goals, through literature refinement, qualification via interviews, and ultimately consolidated into the final outcome specification. The next section presents this information in full.
Figure 1. Outcome specification process.
Source: the Author.
I am grateful to the Italian Red Cross and to the three Food Bank Markets that participated in this study for their willingness to be interviewed and to share relevant documentation
42Food banks are charitable organizations that provide food assistance to individuals and families facing economic hardship, typically through the distribution of donated goods. In recent years, more flexible models have emerged, allowing beneficiaries to select food items rather than receiving standardized parcels. Food bank markets (FBMs) — in Italy, ‘Empori Solidali’ — represent a further development of this model, replicating supermarket-like environments where users can choose food items free of charge.
43The stated goals of the program were extracted from official program documents (email communication to the author) and publicly available materials presenting the markets’ aims and activities (Emporio Solidale Vicenza, n.d.; CRI Giulianova, n.d.; CRI Rieti, n.d.). Across the three FBMs analyzed, these documents articulate a broad set of objectives spanning food provision, social support, and environmental benefits. However, they remain insufficient for outcome appraisal. Goals are formulated in general, often normative terms, without specifying the expected changes in users’ conditions or behaviors. In addition, different analytical levels are conflated, with outcomes, intermediate effects, and program features presented without clear distinction. Indicators, targets, and temporal references are consistently absent. As a result, program goals function primarily as statements of intent rather than as clearly specified and empirically tractable definitions of success.
44The next sections examine the four outcomes identified for FBMs: food security, food waste, broader well-being, and social inclusion. For each outcome, the sections analyze how it is articulated in the program documents, how it is further qualified through document review and interviews, and how intermediate outcomes emerge from this process.
45Official goals. In the program documents, food security is primarily expressed as the provision of food through a self-service model that allows users to select items according to their preferences, moving beyond standardized food parcels. This formulation sits between an output (food provision) and a preliminary outcome (improved access and user choice), without explicit reference to users’ nutritional status or stability of access. Emphasis is placed on program features — particularly point-based systems enabling autonomous selection — rather than on observable changes in user conditions. As a result, food security is framed more as an operational goal than as a goal with clearly defined outcome constructs.
46Outcome specification. Research on food banks has primarily examined their capacity to improve users’ food security, generally pointing to limited or inconsistent effects. Existing studies show that food bank use often coexists with persistent users’ food insecurity, with only marginal improvements over time and limited impact on dietary adequacy (Loopstra & Tarasuk, 2012; Bazerghi et al., 2016; Rizvi et al., 2021). Evidence on dietary quality is mixed: while some studies find relatively better access to fruit and vegetables among food bank users (Bertmann et al., 2021), limitations in the supply and management of fresh produce remain significant (Sengul Orgut et al., 2016). These studies typically use standard measures of household food insecurity related to food intake (e.g., skipping meals, having a balanced diet, being worried about food).
47Evidence from the exploratory interviews indicates that the food aid provided typically covers only a fraction of users’ needs in both quantity and variety and is distributed at relatively long intervals (often monthly). Further, the FBM is designed to provide short-term relief rather than produce sustained improvements. This means that stable changes in food security often reflect broader shifts in individual circumstances (e.g., employment) rather than the direct effects of food aid. This pattern is further reinforced by the presence of long-term users, whose structural conditions (e.g., retired people unable to afford rent and bills) limit their exit from the program, thereby weakening the potential to observe changes in food security outcomes.
48These limitations indicate that, although agreed standard measures of food security appear to be the most intuitive constructs for appraising FBM success, they may be insufficient to capture program effects. This illustrates a construct selection problem: while normatively central, food security is only partially sensitive to the intervention and difficult to attribute due to contextual influences. This may suggest retaining these measures as descriptive for user status, while shifting the focus to the appraisal of more proximal outcomes. In this study, two intermediate dimensions emerged as particularly relevant: the quality of food procurement and changes in users’ behavior.
49A) Food procurement capacity. Food banks in Europe draw the bulk of their stock from donated goods, the EU program for food aid (FEAD), and community food drives (FEBA, 2020; Caritas Italiana, 2018) (see also anonymized reference). Each stream has specific strengths and weaknesses in terms of supply stability, food variety, and shelf life, and none is sufficient on its own to sustain a ‘market-style’ supply of food. These three streams typically need to be integrated, requiring active coordination and resource mobilization, hence placing notable burdens on the managing association. In this respect, this capacity for market-like procurement, rather than a (more modest) food-parcel supply, is not simply an output of the program; it can be understood as an intermediate outcome that impacts food availability and is then included in the appraisal of a program’s success.
50B) Users’ behavior. Users are not passive recipients; their behavior mediates whether access to the FBM translates into improved food security. Since FBM shopping rarely fulfills total dietary needs, beneficiaries must decide if and how to integrate it with other purchases.
51One user response regards budget reallocation: by obtaining part of their food for free, users may redirect saved money toward other essential expenses (e.g., bills, rent). While this may improve overall household welfare, it can obscure effects on food security if the additional resources are not spent on food. A second response involves dietary changes. Achieving a balanced diet requires users to evaluate what to shop for at the FBM and what to buy elsewhere. Given the limited range of items at the FBM, users must rely on it for certain foods and source complementary items elsewhere (e.g., fresh produce, protein sources). This, however, requires both adequate purchasing power and dietary awareness.
52In both cases, the benefits derived from accessing the FBM depend on users’ responses which must be examined to fully capture what specific changes the FBM is producing and for whom, since the capacity for congruent budgetary and dietary responses will not be equally distributed across beneficiaries.
53Official goals. In the program documents, food waste is framed as the recovery and redistribution of surplus food from retailers and local producers. It is presented both as a distinctive feature of the FBM model and as an environmental benefit, contributing to the model’s legitimacy. However, it remains unclear whether it is intended as a program goal or simply as an operational characteristic. The formulation does not refer to user-level changes, and while it implicitly points to societal benefits, no indicators — either of outputs or outcomes — are specified. As a result, food waste prevention functions more as a broad rationale for the program than as a clearly defined outcome.
54Outcome specification. Most food delivered through food banks comes from the recovery and reuse of surplus food. Indeed, food waste reduction and food security are closely linked objectives in most countries, thanks to government schemes that promote the donation of surplus food (Busetti & Pace, 2022; Aitken et al., 2024; Bech-Larsen et al., 2019; Lambie-Mumford et al., 2020). In their evaluation of an Italian FBM, for instance, Ranuzzini & Gallo (2020) developed a cost-benefit model to assess the wider social benefits of the FBM; interestingly, they found that food recovery from waste was necessary to achieve a positive benefit-cost ratio.
55However, the policy link between food waste and food security has been widely criticized for legitimizing a gap in public policy, subordinating the right to food to corporate philanthropy, and perpetuating the continuous overproduction of food by corporate firms (Riches, 2018; Arcuri, 2019). Moreover, donated food may not align with healthy dietary habits or may include items that do not typically enter the market (Busetti & Pace, 2022).
56Evidence from the exploratory interviews further suggests that FBM staff do not consider food waste reduction a relevant outcome. Operating primarily with surplus food donations, FBMs automatically reduce food waste, but staff consistently framed it as an operational necessity driven by donation logistics and supply constraints rather than as an intended outcome of the intervention. Interviews also highlighted that surplus food requires substantial management capacity and, although essential, entails significant operational burdens: it typically has a short shelf life and may need to be ‘promoted’ to users (i.e., excluded from the point system) to avoid becoming waste.
57In this respect, treating food waste reduction as an outcome would misrepresent program effectiveness by considering a consequence of the donation system (i.e., a contextual condition) as a positive program effect. Moreover, it would lead to a paradox in which food banks that rely less on surplus donations — for instance, by purchasing food to improve quality or variety — would appear less effective, even though they might provide better service to users. These considerations illustrate the importance of scrutinizing the outcome domain: not all plausible effects — even when mentioned in official statements — constitute appropriate outcomes for appraising program success. In this case, food waste would not be retained in the outcome domain.
58In the program documents, broader well-being is framed in general terms as the enhancement of users’ overall conditions through connections with local services and voluntary organizations. It is presented as an added value of the FBM model that extends beyond immediate material assistance, but the formulation remains abstract: no specific dimensions of well-being are identified, and no outputs or outcomes are specified. As a result, broader well-being appears as a general program aspiration rather than as a clearly defined, outcome-oriented construct.
59Outcome specification. A distinctive feature of FBMs is the provision of services alongside food aid, possibly including education and training, cultural and recreational initiatives, job counseling, and volunteering opportunities (Sforzi et al., 2022). This orientation has been further reinforced at the policy level by the integration of the Fund for European Aid to the Most Deprived (FEAD) into the broader European Social Fund+ (EFS+), which encourages charitable organizations to offer a wider range of services. In practice, interviews confirm that — although with considerable variation across cases — most FBMs provide at least basic counseling services and may also help identify needs, offer targeted services beyond food aid, and contribute to users’ broader well-being.
60However, the literature on this outcome remains mixed. While some studies report positive effects of integrated models combining food provision with additional services (Martin et al., 2013), evidence on broader socio-economic indicators is limited (Ranuzzini & Gallo, 2020).
61In appraising programmatic success, broader well-being is a problematic outcome. On the one hand, it can be plausibly included in the outcome domain, particularly in contexts of long-term deprivation where integrated forms of support are expected to complement food aid. However, its relevance as a construct depends on program configuration: only some FBMs provide additional services beyond food distribution, making broader well-being a more plausible outcome in those cases than in others. At the same time, the multidimensional nature of deprivation is largely shaped by structural factors — such as income, employment, and housing — that lie beyond the scope of FBMs, limiting the extent to which changes can be attributed to the program. Moreover, changes in broader well-being are likely to be slow-moving, so that short-term evaluations may underestimate impact. Taken together, these features suggest that, while broader well-being can be included in the outcome domain, it does not systematically constitute an analytically suitable construct for capturing program success.
62This perspective is consistent with FBMs’ operational role: rather than directly addressing structural determinants of deprivation, they operate through intermediate outcomes, such as identifying needs, offering complementary services, and connecting users to external support networks. These dimensions are operationalized through the three intermediate outcomes discussed below.
63A) Needs detection. Needs detection refers to FBMs' capacity to identify users’ needs beyond food insecurity. According to the exploratory interviews, FBMs may favor closer interactions between users and staff compared to traditional food parcel delivery, thereby facilitating the identification of a broader range of needs (e.g., health, employment, or social vulnerabilities).
64This capacity depends on establishing a relationship of trust with users and can be strengthened by organizational features that encourage more frequent contact and proximity to the FBM staff: check-in procedures, staff assistance during shopping, and frequency of attendance. Assessing the implementation of such organizational practices may provide essential clues about the plausibility of FBMs’ contribution to users’ broader well-being. However, interview evidence suggests that these practices remain highly dependent on informal, context-specific arrangements rather than standardized procedures.
65B) Service scale-up. Service scale-up refers to FBMs’ capacity to host a variety of additional services beyond food aid. Unlike user-level outcomes, this dimension captures changes in organizational capacity, reflecting the extent to which the intervention enables a diversified provision of support. According to the exploratory interviews, opening the FBM was considered an incentive for expanding the association’s activities. With respect to food parcel storage, having a stable, equipped venue reduces logistical barriers to designing and implementing new initiatives and may also allow the managing association to apply for funding requiring dedicated physical spaces.
66In outcome terms, service scale-up constitutes an intermediate outcome at the organizational level, reflecting changes that precede potential effects on users’ well-being. More broadly, it shows that intermediate outcomes may involve shifts in organizational capacity, which in turn expand the range of outcomes a program can plausibly generate.
67C) Support networking. Support networking refers to FBMs' ability to connect users with external sources of support. Interviews suggested that the FBM may operate as a broker within informal support networks involving other associations and public bodies. However, the extent to which this function is effectively realized depends on local conditions, particularly the capacity to build and maintain partnerships and the willingness of external actors to respond.
In outcome terms, support networking should capture the program’s capacity to activate and mediate such connections. At the same time, the extent to which these connections translate into effective support depends on the response of other organizations and public bodies. As a result, support networking constitutes an outcome that is only partially attributable to the FBM and should be interpreted with caution in relation to broader well-being effects.
68Official goals. In the program documents, social inclusion is framed by elements such as dignity, autonomy, empowerment, and connections to community networks. It is associated with non-stigmatizing modalities of assistance — such as self-service models — as well as with relational forms of support, including listening and counseling. While these elements suggest a broader ambition to address users’ social vulnerability, the formulation remains conceptually ambiguous: multiple dimensions are evoked without a clear indication of intended outcomes or program features. No specific user-level changes are defined, and no measures are specified. As a result, social inclusion qualifies as a broad program objective that is difficult to disentangle and operationalize.
69Outcome specification. Social inclusion refers to the extent to which FBMs enable users to engage in social interactions and build or maintain social relationships. Unlike other outcomes, it does not primarily operate through the provision of material resources or services, but through the social experience associated with participation in the FBM. Social inclusion was included in the outcome domain both with reference to official documents highlighting the importance of dignity, autonomy and comfort in receiving food aid, and following interviews stressing the significance of the FBM in offering a richer social experience, where individuals may find a place to socialize outside their homes.
70This socialization could simply involve talking with volunteers but might also include interactions with other users. Even the simple opportunity to leave their homes could provide a meaningful experience for people who are likely to feel lonely, have fewer relational contacts with friends and family, or risk home isolation due to mobility problems. International evidence also highlights the importance of the social dimension of food aid, with users valuing opportunities for informal interaction and a welcoming environment (Mulrooney et al., 2023). In this respect, social inclusion may be part of the outcome domain and constructed as an increase in users’ social interactions and contacts.
71During interviews, one of the main concerns regarding the measurement of this outcome was the need for a longer timeframe for socialization to unfold. Users may initially be reluctant to engage, even when prompted by minimal forms of activation, such as attending the FBM instead of quickly collecting a standard food parcel. Overall, socialization typically has a threshold dynamic: rapport builds after multiple visits and at higher visit frequency. This highlights the timing dimension of outcome appraisal: although social inclusion is a relevant construct, its empirical observability may depend heavily on temporal dynamics and may be impeded by short-term assessments.
72Drawing on these preliminary data, two intermediate outcomes emerged as relevant: A) a normalized shopping experience, and B) a perception of comfort by users.
73A) Normalized shopping experience. A normalized shopping experience refers to the extent to which FBMs reduce the stigma associated with receiving food aid by replicating features of ordinary retail environments. Key design features of the FBM shape the degree of stigma associated with being a user: the physical layout, the food on offer, access, and shopping procedures.
74Physical layouts that mimic an ordinary supermarket may reduce the visibility of being an aid recipient, thereby encouraging regular attendance and interaction. Similarly, the availability of familiar brands rather than aid-branded food, together with a market-like shelf layout, may contribute to a sense of normalcy and dignity. Finally, access and shopping procedures may also directly structure opportunities for interaction, although these design features are constrained by reliance on volunteers. As suggested by the interviews, associations often limit opening hours and regulate access to manage demand, in most cases allowing access only by appointment (see also Saxena & Tornaghi, 2018). In this respect, opportunities for social interaction depend only in part on the degree of normalization in product offerings and physical layouts; they are primarily shaped by organizational arrangements regarding access, which determine where and how contact occurs (e.g., during waiting times, inside the market, with other users or mainly with volunteers).
75Given these constraints, the extent to which a normalized shopping experience is achieved — and how associations cope with such limitations — should be explicitly assessed, as it may act as a key moderator of social inclusion effects.
76B) Perceived stigma and user comfort. Perceived stigma and user comfort refer to users’ subjective experience of accessing the FBM, particularly the extent to which participation is perceived as socially acceptable and non-stigmatizing.
77Even when the FBM replicates an ordinary shopping environment, receiving charitable aid may still generate discomfort. Indeed, users may differ in their sensitivity to stigma, even within the same environment. Interviews suggested that this sensitivity is particularly acute among individuals who are not accustomed to using social services, such as those who recently lost their jobs and see their need for assistance as temporary. Feelings of shame or of “not belonging” to the aid-recipient group can limit both the frequency and the quality of engagement with the FBM, reducing its potential as a social space.
78These considerations suggest that user variation in perceived stigma and comfort may condition patterns of participation, thereby influencing the extent to which social inclusion can develop across user groups.
79The empirical analysis provides a set of insights that refine and extend the four analytical challenges identified in the appraisal of programmatic success, showing how outcome theorization operates and how it reshapes what counts as program success.
80A first implication concerns the specification of the outcome domain. The analysis shows that official program goals do not translate directly into relevant outcomes. In the case of FBMs, program documents articulate a broad range of aims that vary significantly in their analytical status. While some can be readily translated into outcomes, others are formulated as general orientations or operational features and require additional analytical work to be part of the outcome domain (e.g., from ‘relational support’ to ‘social inclusion’). At the same time, some objectives — such as food waste reduction — are more appropriately understood as features of the supply system, reflecting the use of surplus food within existing donation schemes, rather than as program effects, and therefore excluded from the outcome domain on analytical rather than normative grounds.
81These results suggest that outcome specification does not simply require adding to official goals but also entails excluding misleading effects and transforming vague aims into theoretically grounded, observable outcomes. While the theoretical discussion stressed the potential relevance of unintended effects, the empirical analysis provides only limited evidence of their role in this case, suggesting caution in assessing how well a theory-driven approach can anticipate them during ex ante outcome specification.
82A second implication relates to construct selection. The case highlights a tension between the centrality of certain outcomes and their analytical suitability for capturing program effects. Food security, for instance, represents the most direct and policy-relevant outcome of food aid, yet the empirical analysis shows that widely accepted standard measures are only partially sensitive to FBM interventions and strongly influenced by exogenous factors. As a result, they may yield limited or ambiguous evidence of program effects - increased food availability due to the FBM shopping may coexist with persistent food insecurity in terms of standard measures of food intake. A similar issue emerges with broader well-being: while often associated with FBMs in program narratives, its realization depends on the provision of additional services and on the effectiveness of external actors, such as local authorities or partner organizations. These actors lie beyond the direct control of the program.
83These examples confirm that selecting outcome constructs cannot be based solely on their relevance to the policy problem or program goals, but must be grounded in their alignment with the program’s causal reach and their capacity to detect change. In this respect, the analysis suggests that normatively central outcomes may not always constitute analytically appropriate constructs of program success, and that alternative outcomes — often more proximal — may provide a more accurate representation of program effects.
84A third implication concerns the role of intermediate outcomes. The findings show that program effects are not only reflected in distal outcomes but in a set of intermediate changes that reflect how the program operates. In the case of FBMs, these include organizational capacities such as food procurement and service scale-up, as well as user-level processes such as behavioral adaptation, needs detection, and engagement with support networks. These dimensions are not merely transitional steps towards final outcomes, but constitute core components of program performance, as they directly condition the program’s ability to generate broader effects. This expands the role typically attributed to intermediate outcomes, suggesting that they should be treated not only as explanatory elements within causal models, but also as relevant dimensions of programmatic success in their own right.
85A fourth implication relates to the temporal dimension of outcomes. The analysis shows that different outcomes follow distinct temporal dynamics that affect their observability and interpretation. In the FBM case, changes in food security tend to be slow and heavily conditioned by structural factors. By contrast, social inclusion develops gradually through repeated participation and interaction, often requiring sustained engagement before becoming observable. Intermediate outcomes, such as procurement capacity or needs detection, tend instead to emerge over shorter timeframes, reflecting more immediate program effects. These differences confirm that outcome appraisal depends not only on what is measured, but also on when it is measured. Without explicit consideration of these temporal dynamics, appraisals of program success risk underestimating effects, overlooking transient changes, or misinterpreting variation over time as failure or success.
86The analysis developed in this article does not concern whether policies succeed or fail, nor whether the criteria of success are contested among stakeholders, questions that the policy success literature has addressed extensively. It concerns what happens analytically before any of that: the consequential choices that determine which outcomes are included, how they are conceptualized, where they are located in the causal chain, and when they are expected to emerge. These choices are not peripheral to the appraisal of programmatic success, but are constitutive of it and have remained largely implicit in a literature that has otherwise grown considerably in the appraisal of multiple dimensions of policy success.
87The core argument is that programmatic success is not observed but constructed, not in terms of the social construction of the political narrative of success, but as the analytical construction of outcomes as objects of evaluation. The article’s first contribution lies in repositioning outcome specification within the policy success framework. The programmatic dimension has been treated as the least theoretically challenging component of the framework, with outcomes assumed as derivable from goals or problem definitions. The analysis shows this assumption is unwarranted: different outcome configurations yield different assessments of program performance, and the choices involved are consequential enough to determine whether a program appears to succeed or fail. Addressing this requires recognizing that the programmatic dimension presents a prior analytical problem that has not yet been systematically addressed, one that is preliminary to measurement and causal attribution.
88A second contribution lies in reorienting theory-driven evaluation toward this problem. TDE has primarily been used to explain how and why outcomes occur - to open the black box between program activities and effects. This article shows that TDE reasoning can be mobilized earlier, at the stage of outcome theorization, before causal modeling and measurement begin. In this respect, theory not only accounts for outcomes but defines which outcomes are plausible, analytically appropriate, and observable. In this sense, the article recovers an insight present in the earliest formulations of theory-driven evaluation - that theorizing outcomes is a precondition for evaluating them.
89A third implication concerns the relationship between outcome theorization and program design. If different outcome configurations produce different appraisals of the same program, then the choice of outcomes is not merely an evaluative decision; it may support design choices by highlighting what the program is expected to do and what counts as working. This has practical consequences. Intermediate outcomes, in particular, are closer to program activities and are more directly shaped by design features: access rules, service configurations, organizational procedures. Making them explicit as dimensions of program appraisal does not simply improve evaluation; it identifies where design can intervene. In the FBM case, for instance, whether social inclusion is a plausible outcome depends directly on organizational arrangements that structure the conditions for user interaction. These findings also speak to the growing ambition to reposition policy analysis as a design-oriented field. As Compton et al. (2019, p. 120) argue, advancing policy design requires "a richer and more relevant conceptualization of its key dependent variable" - a condition that outcome theorization directly addresses.
90These contributions are necessarily circumscribed by the scope of the analysis. Outcome theorization as developed here is a preliminary step, not a substitute for evaluation. The FBM case illustrates the approach in a specific organizational and institutional context, and the limited number of interviews constrains the range of empirical variation captured. More fundamentally, the approach relies on existing theoretical knowledge and exploratory evidence to generate outcome hypotheses - knowledge that may be scarce for novel or poorly studied interventions and that must ultimately be tested through rigorous empirical analysis. Yet, in the absence of explicit theorization, analysts risk defaulting to standardized outcomes that reflect the policy problem in general rather than the specific effects of the program under study. What the approach can claim is not predictive completeness but analytical discipline: it makes the assumptions underlying outcome choice explicit and examinable, reducing the risk that appraisals of programmatic success rest on constructs that are convenient but poorly suited to what programs actually do. That, in itself, is a necessary condition for learning from policy.