/
Measurement Validity and Program Evaluation
Save to my account
Sign up
Measurement Validity and Program Evaluation
Chapters 12-13 Flashcards
Study
1
Question
How does construct underrepresentation differ from construct irrelevance in measurement development?
Page 2
Answer
Construct underrepresentation means a measure is too narrow, failing to include key facets of the target construct. Construct irrelevance means a measure is polluted by extraneous variables that distort test scores. Hook: Picture Underrepresentation as an undersized net letting crucial fish escape, whereas Irrelevance is dragging up irrelevant trash.
2
Question
What are the five essential types of validity evidence identified by AERA et al. (2014)?
Page 2
Answer
Test content, response processes, internal structure, relations to other variables, and consequences of testing. Hook: Remember the mnemonic acronym 'CRISC' (Content, Response, Internal, Relations, Consequences).
3
Question
How do researchers empirically gather evidence based on test content during instrument design?
Page 3
Answer
Researchers review existing literature to align items with a theoretical framework and consult panels of subject-matter experts to review item relevance and clarity.
4
Question
Why is exploratory factor analysis (EFA) utilized before confirmatory factor analysis (CFA)?
Page 3
Answer
EFA discovers the underlying latent structure among items without a rigid prior model. CFA formally tests whether empirical data fit a pre-specified, theory-driven factor structure. Hook: EFA is an Explorer scouting unknown territory; CFA is a Judge confirming the established map.
5
Question
What memory hook represents the contrast between exploratory factor analysis and confirmatory factor analysis?
Page 3
Answer
EFA is an Explorer scouting unmapped terrain, whereas CFA is a Judge Confirming an established boundary map.
6
Question
What sample size ratio is required when conducting exploratory or confirmatory factor analysis?
Page 3
Answer
A minimum of 10 to 20 participants per item is required. Hook: Imagine 10 to 20 judges lined up behind every single survey item to evaluate it.
7
Question
Why must researchers collect separate samples when conducting both EFA and CFA in a single study?
Page 3
Answer
Running CFA on the same dataset used for EFA creates circular capitalization on chance. A separate sample cross-validates the discovered factor structure on independent data. Hook: Double the analyses, double the respondents.
8
Question
Why does a statistically significant chi-square test often occur in structural equation modeling despite acceptable fit indices?
Page 3
Answer
Chi-square statistics are hypersensitive to large sample sizes, causing trivial discrepancies between observed and hypothesized models to reach statistical significance. Alternative indices like CFI and TLI are consulted instead.
9
Question
Why is reliability classified as a property of scores rather than a property of the instrument?
Page 4
Answer
Reliability varies across different populations, settings, and administrations. A scale cannot possess static inherent consistency across all possible participant groups.
10
Question
What memory hook distinguishes construct underrepresentation from construct irrelevance?
Page 2
Answer
Underrepresentation is an undersized net missing the target fish; Irrelevance is pulling up tangled, extraneous seaweed along with the catch.
11
Question
How does McDonald's omega provide a theoretical advantage over Cronbach's alpha for estimating internal consistency?
Page 4
Answer
McDonald's omega relies on fewer restrictive mathematical assumptions, such as tau-equivalence, yielding a more accurate reflection of composite reliability.
12
Question
How do convergent evidence and discriminant evidence function collaboratively to validate a new construct scale?
Page 4
Answer
Convergent evidence confirms strong statistical correlations with established instruments measuring related constructs. Discriminant evidence confirms weak or non-significant correlations with instruments measuring unrelated constructs.
13
Question
How does program evaluation differ fundamentally from research under the federal OHRP definition?
Page 6
Answer
Research seeks systematically to generate generalizable knowledge across broader populations. Program evaluation focuses specifically on the merit, worth, or decision-making needs of a specific local program.
14
Question
How does summative evaluation differ in function and focus from formative evaluation?
Page 7
Answer
Summative evaluation assesses overall bottom-line efficacy and outcomes to decide continuation. Formative evaluation monitors ongoing implementation and process mechanics to guide real-time program refinement.
15
Question
What memory hook captures the functional difference between formative evaluation and summative evaluation?
Page 7
Answer
Formative evaluation is the chef tasting the soup during cooking; Summative evaluation is the dining guest eating and judging the finished dish.
16
Question
Why did nineteenth-century 'payment by results' educational evaluation systems decline across England and the United States?
Page 8
Answer
Funding tied strictly to test performance prompted teachers to narrow instruction solely toward exam preparation rather than deep learning, proving inefficient and distorting pedagogical quality.
17
Question
How did the Tylerian Age transform educational and program evaluation methodologies during the 1930s?
Page 8
Answer
Ralph Tyler linked program objectives directly to observable client behaviors, establishing the foundation for criterion-referenced evaluation rather than subjective impressions.
18
Question
How did Bloom's Taxonomy influence program evaluation theory during the Age of Innocence?
Page 8
Answer
It framed student cognition as complex and multidimensional, allowing evaluators to classify learning outcomes hierarchically rather than treating participants as uniform, single-variable units.
19
Question
Which two historical geopolitical events catalyzed federal evaluation mandates during the Age of Development?
Page 9
Answer
The Soviet launch of Sputnik in 1957, which triggered the National Defense Education Act, and the passage of the Elementary and Secondary Education Act in 1965.
20
Question
What memory hook links the chronological historical eras of program evaluation from 1792 to the present?
Page 8
Answer
Remember 'Red Elephants Take Innocent Dogs Past Earth' for Reform (1792), Efficiency (1900), Tylerian (1930), Innocence (1946), Development (1959), Professionalization (1973), and Expansion (1983).
21
Question
What five standards did Stufflebeam et al. (2000) identify during the Age of Professionalization?
Page 9
Answer
Evaluations must serve client needs, address central values, handle situational realities, fulfill probity requirements, and satisfy veracity needs.
22
Question
What four basic cultural competence reminders did the CDC highlight for program evaluators?
Page 9
Answer
No evaluation is culture-free; competence requires self-reflection; competence in one context does not guarantee competence elsewhere; and cultural competence builds essential stakeholder trust.
23
Question
What sequence defines the first three sequential steps in the CDC program evaluation framework?
Page 10
Answer
Step 1 involves identifying and engaging stakeholders; Step 2 defines the program's intended mission; Step 3 focuses the design of the evaluation plan.
24
Question
What sequence defines steps four through six in the CDC program evaluation framework?
Page 11
Answer
Step 4 is gathering credible evidence; Step 5 is analyzing data and drawing conclusions; Step 6 is presenting findings and ensuring intentional utilization.
25
Question
What memory hook represents the full six-step CDC program evaluation sequence?
Page 10
Answer
Remember 'Smart Dogs Dig Excellent Artifacts Promptly' for Stakeholders, Define program, Design focus, Evidence gathering, Analyze conclusions, and Present/use findings.
26
Question
Why is Step 6 (Present Your Findings and Use Them) identified as the step that most often falls short?
Page 11
Answer
Programs frequently focus all institutional energy on the logistics of data gathering, leaving results unused in filed reports rather than driving intentional administrative improvement.
27
Question
How does a logic model depict the structural theory of an intervention?
Page 11
Answer
It provides a clear graphical blueprint mapping investments to real-world results: Inputs funnel into concrete Activities and Outputs, which systematically generate Short-, Medium-, and Long-term Outcomes.
28
Question
How do program outputs differ conceptually from program outcomes within a logic model framework?
Page 11
Answer
Outputs represent completed programmatic activities and participation counts. Outcomes represent actual substantive shifts in client knowledge, attitudes, behaviors, or health status resulting from those activities.
29
Question
What concrete operational items constitute inputs in a counseling group logic model?
Page 11
Answer
Inputs encompass quantifiable operating assets, including counselor hours, allocated office square footage, utility costs, preparation time, and client curriculum materials.
30
Question
How does person-centered planning enrich program evaluation data collection in community mental health?
Page 7
Answer
It directly solicits client and family definitions of recovery, surfacing unpredicted subjective priorities that standard clinical metrics overlook and aligning evaluation with authentic client need.