Analysis of Internal Conflict

School of Social, Political and International Studies
Universidad del Rosario

Session 2:2 · Basic Quantitative Analysis

The Recruitment of Children and Teenagers into the Colombian Conflict

Causality, association and evidence

Mathew H. Charles, PhD

01

Begin with the question

What are we trying to explain?

Is the recruitment of children and teenagers more closely associated with:

A

Structural social conditions

Poverty, school non-attendance, educational delay, child labour and economic dependency.

or
B

Conflict dynamics

Armed-group presence and rivalry, displacement, coca cultivation, targeted murders and territorial violence.

Research purpose: compare patterns, identify plausible explanations and learn how far the evidence allows us to go.
02

From question to expectation

What is our hypothesis?

A hypothesis is simply the answer we expect before looking at the results. The analysis then checks whether the observed pattern supports it.

Main hypothesis

Conflict dynamics matter more

Municipal differences in recruitment will be more strongly related to armed-group presence, territorial competition, displacement and conflict-related violence than to poverty and educational disadvantage alone.

What patterns would support it?
  • Higher recruitment where two or more armed-actor families are present.
  • Higher recruitment in territories with rearmed-FARC presence, displacement, coca cultivation or attacks on social leaders.
  • Conflict indicators showing stronger correlations than social indicators.
What the dataset cannot test directly

It records whether groups are present in the same municipality. It does not record whether they were fighting each other. Co-presence can therefore suggest possible rivalry, but it is not the same as a direct rivalry measure.

03

Descriptive statistics

First describe the evidence universe

Descriptive statistics tell us what the dataset contains before we test relationships.

1,120municipalitiesone row per municipality
828documented cases2017–2020 total
143municipalities with cases12.8% of municipalities
1,100municipal rates available20 lack a complete denominator
977municipalities record zero cases
143municipalities record at least one case

Important: zero means “no documented case in this dataset,” not “recruitment did not occur.”

04

Spatial description

Recruitment is territorially concentrated

Maps reveal clusters, but they do not explain why those clusters exist.

Map and table of recruitment cases by Colombian department, 2017 to 2020
Department totals: useful for identifying broad regional concentration.
Map of Colombian municipalities with the largest recruitment counts, 2017 to 2020
Municipal concentrations: the national pattern is produced by particular local hotspots.

Original descriptive graphics supplied with the dataset; integrated here as evidence, not treated as a causal model.

05

Raw counts

Which departments contain the most documented cases?

Counts identify where the largest number of cases was documented.

Cordoba
172
Antioquia
142
Choco
99.0
Nariño
51.0
Caqueta
48.0
Vaupes
42.0
Guaviare
41.0
Cauca
37.0
Arauca
36.0
Norte De Santander
33.0
Bolivar
23.0
Meta
20.0

Do not confuse volume with individual risk: a larger department or youth population can generate more cases even when its rate is lower.

06

Local concentration

A small group of municipalities shapes the national total

Puerto Libertador
58.0
Montelibano
43.0
Caceres
42.0
Monteria
37.0
San Jose De Ure
30.0
Bajo Baudo
25.0
Tado
23.0
San Jose Del Guaviare
23.0
San Andres De Tumaco
23.0
Medellin
19.0
MunicipalityCasesRate per 10,000
01Puerto LibertadorCordoba
5831.7
02MontelibanoCordoba
4313.0
03CaceresAntioquia
4231.1
04MonteriaCordoba
372.2
05San Jose De UreCordoba
3048.7
06Bajo BaudoChoco
2516.8
07TadoChoco
2330.4
08San Jose Del GuaviareGuaviare
2310.2
09San Andres De TumacoNariño
232.3
10MedellinAntioquia
190.3
07

Choosing the outcome

Why calculate recruitment per 10,000?

A rate makes municipalities with different youth populations more comparable.

Recruitment rate
documented recruitment cases÷population aged 0–19×10,000

The answer estimates how many documented cases there are for every 10,000 young people.

Municipality AMore cases
50cases÷100,000young people
5 per 10,000
Municipality BHigher rate
20cases÷10,000young people
20 per 10,000

Interpretation: Municipality A has more cases, but Municipality B has the higher recruitment rate relative to its youth population.

08

From one municipality to a group

What is the mean recruitment rate?

Mean simply means average. Imagine that three municipalities belong to the same armed-actor category.

Municipality A1case per 10,000
Municipality B3cases per 10,000
Municipality C5cases per 10,000
Step 1 · Add the three rates1 + 3 + 5 = 9
Step 2 · Divide by three municipalities9 ÷ 3 = 3
The group mean3 per 10,000

What this tells us: 3 is one summary number for the group.

What it does not tell us: Municipality A still has a rate of 1 and Municipality C still has a rate of 5.

09

Two ways to summarize a group

What is the difference between mean and median?

We use both because recruitment rates contain many zeros and a small number of very high values.

Five municipal rates, placed in order
000middle218very high
Mean = average(0 + 0 + 0 + 2 + 18) ÷ 5 = 4

The very high value of 18 pulls the mean upward.

Median = middle value0

After ordering the five rates, the third—or middle—value is 0.

Mean: uses every value and captures the overall level, but can be pulled upward by a few extreme municipalities.

Median: describes the middle municipality and is less affected by extremes, but can remain zero when fewer than half of municipalities record recruitment.

How to read the later table: when the mean is much higher than the median, a few high-rate municipalities are raising the average. That is why showing both is useful.

10

Conflict dynamics

Recruitment rises sharply where armed actors overlap

The bars show the mean of the available municipal recruitment rates; the percentages show how many municipalities record any recruitment.

What the bars compare:the average municipal rate in each actor-presence category. They do not claim that every municipality in a category has that rate.
11

Composition matters

Not every form of co-presence produces the same pattern

Mean shows the average rate; median shows the middle municipality. A large gap warns us that a few high-rate municipalities are pulling up the average.

Recorded actor combinationMunicipalitiesRecruitment-positiveCasesMean rate
(average)
Median rate
(middle)
Dissident FARC + Paramilitary successor4351.2%2175.090.49
Dissident FARC4438.6%1134.380.00
ELN + Paramilitary successor4362.8%1553.680.74
Dissident FARC + ELN1361.5%622.631.83
Dissident FARC + ELN + Paramilitary successor3271.9%1002.371.28
ELN4012.5%180.610.00
Paramilitary successor21810.1%1120.230.00
No recorded actor family6872.8%510.120.00

Co-presence is a proxy: it tells us that actor families overlap territorially. It does not tell us whether they were actively fighting, cooperating or operating in different parts of the municipality.

12

Rearmed FARC

A strong territorial association—with an important warning

No dissident category recorded0.36mean cases per 10,000
7.8% recruitment-positive

371 cases across 1012 municipalities

Residual FARC only1.82mean cases per 10,000
52.5% recruitment-positive

119 cases across 59 municipalities

Rearmed FARC recorded7.70mean cases per 10,000
67.3% recruitment-positive

338 cases across 49 municipalities

Presence is not perpetration.

Municipalities recording rearmed-FARC presence contain 338 cases, but the separate attribution file assigns many cases in those territories to the AGC or leaves the actor unknown. Rearmed-FARC presence may be marking contested territories rather than recruitment committed by one actor.

13

Reported perpetrator

Can we say that one group recruits more than another?

Only cautiously: 43.3% of the actor-attribution file is coded as unknown.

Unknown actor
354
AGC
174
Dissident FARC
136
ELN
103
Other actors
48.0
354cases coded “unknown actor”

The largest category is missing attribution. Therefore, the ranking of identified actors is incomplete.

Data audit: this file contains 818 recruitment cases and its actor columns sum to 815, while the enhanced universe contains 828. Use the actor ranking as preliminary descriptive evidence until the source versions are reconciled.

14

Before the scatterplot

What is a correlation coefficient?

A correlation coefficient is one number that summarizes how two variables vary together across municipalities.

A variable is something that can differ from one municipality to another—for example, recruitment rate, displacement or poverty.

The coefficient asks: when one variable is higher, does the other usually tend to be higher, lower or show no clear pattern?

+0.54
−1strong negative0no clear pattern+1strong positive
First: read the sign

Which direction?

+ means higher values of one variable tend to accompany higher values of the other. means higher values of one tend to accompany lower values of the other.

Second: ignore the sign briefly

How strong?

Look at the distance from 0. A number near 0 is weak. A number nearer 1—positive or negative—is stronger.

Our example

rₛ = +0.54

This means a moderate positive association. It is not 54%, and it does not say how many extra cases one factor produces.

What the symbol means: rₛ is the Spearman correlation coefficient. It describes an association, not causation.

15

Now see the pattern

How to read a scatterplot

The coefficient summarizes the relationship in one number. The scatterplot shows the municipal observations behind that number.

Possible explanatory variable (X)Recruitment rate (Y)positive directionoutlier
  1. Read the axes.X is the possible explanatory variable; Y is recruitment per 10,000.
  2. Each dot is one municipality.Dots close together have similar values.
  3. Look for direction.Upward = positive; downward = negative; no direction = weak or no association.
  4. Judge strength.A narrow cloud around a line is stronger than a widely dispersed cloud.
  5. Notice outliers and zeros.They may reveal exceptional cases, data problems or different mechanisms.
16

Conflict dynamics

Explore the conflict-dynamics relationships

Choose one variable at a time. Recruitment per 10,000 always remains on the vertical axis so the graphs can be compared.

Spearman correlationrₛ = 0.47
No zoom needed.This variable has only a few fixed categories, so hiding the highest 2% would not change the graph.
0.0020.240.360.580.61010.000.601.201.802.403.00Number of armed-actor families (actor families)Recruitment per 10,000 aged 0–19
What the dots show

Number of armed-actor families

The dots form four vertical stacks because the X variable can only be 0, 1, 2 or 3. Recruitment-positive municipalities rise from 2% with no recorded actor family to 15% with one, 58% with two and 72% with three.

What this supports

A clear positive association between armed-actor overlap and recruitment. The stacks show categories, not a smooth continuous increase.

Coefficient: rₛ = 0.47 (moderate positive)N: 1,100 municipalities
recorded territorial presence. Dots can overlap when municipalities have the same values. The pattern shows association, not causation.

For 0/1 presence indicators: the two vertical bands mean “absent” and “present”; their difference is easier to read as a comparison of group means. Multiple presence measures territorial overlap, not observed rivalry or combat.

17

Social conditions

Explore the structural-social relationships

Look first at the direction and spread of the dots. Do not decide that a relationship is strong simply because the line slopes upward.

Spearman correlationrₛ = 0.24
Full-range view.All X values are visible. A few extremely high values may squeeze most dots together on the left.
0.0020.240.360.580.61010.0018.436.955.373.892.2Multidimensional poverty (%)Recruitment per 10,000 aged 0–19
What the dots show

Multidimensional poverty

Dots with zero recruitment appear across the entire poverty range, so poverty alone does not separate municipalities neatly. Even so, recruitment is documented in 27% of the highest-poverty quarter, compared with 7% of the lowest-poverty quarter.

What this supports

A weak positive association: poverty may contribute to vulnerability, but it does not by itself explain where recruitment is documented.

Coefficient: rₛ = 0.24 (weak positive)N: 1,100 municipalities
2018 structural baseline. Dots can overlap when municipalities have the same values. The pattern shows association, not causation.
18

Illicit economies

Explore the illicit-economy relationships

Coca cultivation and drug seizures are not the same type of indicator. Seizures may reflect trafficking, policing or reporting, so interpret them carefully.

Spearman correlationrₛ = 0.49
Full-range view.All X values are visible. A few extremely high values may squeeze most dots together on the left.
0.0020.240.360.580.61010.003,5277,05410.6k14.1k17.6kCoca cultivation (average hectares)Recruitment per 10,000 aged 0–19
What the dots show

Coca cultivation

Most municipalities record no coca cultivation, forming a dense left-hand stack. Recruitment appears in 46% of coca-growing municipalities versus 5% of non-coca municipalities, while a few very large cultivation values stretch the axis.

What this supports

A moderate positive association, consistent with illicit territorial economies forming part of the conflict environment.

Coefficient: rₛ = 0.49 (moderate positive)N: 1,100 municipalities
2017–2020 municipal average. Dots can overlap when municipalities have the same values. The pattern shows association, not causation.
19

Let the software do the calculation

How do we obtain the coefficient in SPSS?

You do not need to calculate the formula by hand. SPSS calculates it; your job is to choose the correct variables and interpret the output.

AnalyzeCorrelateBivariate
  1. Choose two variables.For example: Multiple armed-group presence and Recruitment rate per 10,000.
  2. Move both variables into the Variables box.SPSS will compare every municipality that has a value for both.
  3. Select Spearman.We use Spearman here because the data contain many zeros and extreme values. It compares the order or rank of municipalities.
  4. Click OK.SPSS produces a table. We then read only three elements: the coefficient, the significance value and N.
For this class: students are not assessed on remembering the SPSS menu or formula. They are assessed on interpreting the graph and the output correctly.
20

Reading the SPSS output

First, find the relationship in the table

The table looks complicated because it repeats information. We only need the cells where the two different variables cross.

Correlations
Multiple armed-group presenceRecruitment rate
Multiple armed-group presenceCorrelation Coefficient1.0000.539**
Sig. (2-tailed)< 0.001
N11001100
Recruitment rateCorrelation Coefficient0.539**1.000
Sig. (2-tailed)< 0.001
N11001100

** Correlation is significant at the 0.01 level (2-tailed).

1

Ignore the 1.000 diagonal

A variable always correlates perfectly with itself. These diagonal cells do not answer our research question.

2

Find where the variables cross

Read the yellow cells: Multiple armed-group presence × Recruitment rate.

3

The result appears twice

+0.539 appears in the top-right and bottom-left because the same relationship is shown in both directions.

4

Read the three lines together

Coefficient, Sig. (2-tailed), and N each answer a different question. The next slide explains them.

21

Reading each element

What do the three lines in the coefficient table mean?

Read from top to bottom. Do not interpret one line without the others.

Line 1

Correlation Coefficient: +0.539

Question answered: What is the direction and strength?

The plus sign means a positive direction. The size, 0.539, indicates a moderate relationship. It is not 53.9% and not a predicted number of cases.

Line 2

Sig. (2-tailed): < 0.001

Question answered: Is the pattern unlikely to be a chance result if there were no relationship?

This is the p-value. Below 0.05 is conventionally called statistically significant. “2-tailed” means SPSS checked for either a positive or a negative relationship.

Line 3

N: 1,100

Question answered: How many observations were compared?

Here N means municipalities with usable values for both variables. It is not the number of recruitment cases. Municipalities missing either value are excluded.

What do ** mean?SPSS adds stars as a shortcut for statistical significance. Stars do not show strength, substantive importance or causation.
22

Worked example 1

A clear, statistically significant relationship

Multiple armed-group presence and recruitment per 10,000.

Correlations
Multiple armed-group presenceRecruitment rate
Multiple armed-group presenceCorrelation Coefficient1.0000.539**
Sig. (2-tailed)< 0.001
N11001100

** Correlation is significant at the 0.01 level (2-tailed).

  1. Direction

    Positive. Municipalities with multiple armed groups tend to have higher recruitment rates.

  2. Strength

    0.539 is moderate. It is not perfect, but it is one of the strongest relationships in this dataset.

  3. Significance

    p < 0.001. This is below 0.05, so we describe the relationship as statistically significant.

  4. Conclusion

    The result supports the conflict-dynamics hypothesis. It does not prove that rivalry caused recruitment.

23

Worked example 2

A weak, non-significant relationship

Heroin seizures and recruitment per 10,000.

Correlations
Heroin seizuresRecruitment rate
Heroin seizuresCorrelation Coefficient1.0000.043
Sig. (2-tailed)0.158
N11001100
  1. Direction

    The sign is positive, but the coefficient is extremely close to zero.

  2. Strength

    0.043 is very weak. The municipalities do not form a clear upward or downward pattern.

  3. Significance

    p = 0.158. This is above 0.05, so the result is not statistically significant.

  4. Conclusion

    This dataset provides no clear evidence of a relationship between heroin seizures and recruitment. That is not the same as proving that no relationship can ever exist.

24

Return to the research question

Which category shows the stronger relationships?

Now that we know how to read a coefficient, we can compare the results. Longer bars mean stronger positive associations with recruitment per 10,000.

Conflict dynamicsMultiple armed-group presence
0.54
Conflict dynamicsCoca cultivation
0.49
Conflict dynamicsNumber of armed-actor families
0.47
Conflict dynamicsForced displacement
0.46
Conflict dynamicsEarly warnings
0.46
Conflict dynamicsDissident-FARC presence
0.46
Conflict dynamicsMurders of social leaders
0.43
School non-attendance
0.33
Educational delay
0.31
Multidimensional poverty
0.24
Child labour
0.22
Unemployment
0.07
Answer:the conflict indicators generally show stronger relationships than the social indicators. This supports our hypothesis, but it does not yet establish causation.
25

Tempering the argument

If correlation is not causation, why bother?

1

Description

Where and how often does recruitment appear?

2

Association

Which factors systematically occur alongside it?

3

Explanation

What plausible mechanism links the factors?

4

Causal inference

Would recruitment have differed without the factor?

Correlation is valuable because it can:

  • rule out simplistic claims that do not fit the pattern;
  • identify municipalities and mechanisms for qualitative investigation;
  • compare the relative strength of competing explanations;
  • generate better causal hypotheses;
  • show where arguments must be qualified.

Appropriate conclusion: “The evidence is consistent with conflict dynamics playing a stronger role.” Not: “The data prove conflict caused recruitment.”

26

From pattern to explanation

Quantitative results help us choose cases for qualitative investigation

The numbers show where an expected pattern appears—and where it does not. Interviews and documents can then investigate how and why.

Conflict dynamicshigher ↑lower ↓
Higher recruitmentLower recruitment
Expected pattern

High conflict
High recruitment

How did conflict exposure lead to recruitment?

Protective puzzle

High conflict
Low recruitment

What interrupted or prevented recruitment?

Alternative puzzle

Low conflict
High recruitment

What explanation are we missing?

Baseline case

Low conflict
Low recruitment

What distinguishes this context?

Quantitative data: identifies the pattern and selects contrasting municipalities.

Qualitative data: interviews, early warnings and local documents trace mechanisms, protection and missing explanations.

27

Strengthening causal claims

How could we get closer to causation?

We cannot force cross-sectional data to prove causation. We improve the research design.

01

Temporal order

Show that rivalry or actor entry occurred before recruitment increased.

02

Annual analysis

Compare changes within the same municipality from 2017 to 2020.

03

Control alternatives

Consider poverty, education, population, coca and displacement together.

04

Matched comparison

Compare similar municipalities that differ in actor competition.

05

Process tracing

Use alerts and interviews to identify how rivalry creates pressure to recruit.

06

Rivalry coding

Distinguish co-presence from explicit territorial contestation.

Mixed-methods bridge: the quantitative analysis identifies the pattern; qualitative evidence investigates the mechanism and tests whether the causal story is credible.

28

Measurement and missing cases

The outcome itself is shaped by reporting systems

Broader documented universe828official records + COALICO monitoring + fieldwork, according to the legacy key
Official Unidad extract217registered victims, 2017 to 1 September 2020
611numerical differencenot proven hidden cases
41municipalities with zero official casesdespite broader documented cases
246cases in those municipalities29.7% of the broader total

Correct language: this indicates an official-registration gap. It is not a definitive estimate of all unreported recruitment.

29

Case comparison

Under-reporting varies dramatically between municipalities

MunicipalityDepartmentBroader countOfficial countNumerical share
Puerto LibertadorCórdoba5800.0%
MontelíbanoCórdoba4300.0%
San José de UréCórdoba3000.0%
CarurúVaupés1800.0%
MitúVaupés1200.0%
SoachaCundinamarca1000.0%
AraucaArauca11763.6%
TameArauca7457.1%
Puerto RicoCaquetá88100.0%
30

Interpretation exercise

What we can and cannot conclude

31

Session synthesis

Five things to retain

  1. 01

    Start with a clear research question and hypotheses.

  2. 02

    Describe the outcome before testing explanations: counts and rates answer different questions.

  3. 03

    Read scatterplots through axes, dots, direction, strength and outliers—not only the coefficient.

  4. 04

    Conflict dynamics show the stronger preliminary associations, especially multiple armed-actor presence.

  5. 05

    Association narrows the argument; causal claims require time, comparison, controls and qualitative evidence.

Evidence should make our arguments more precise—not more certain than the data allow.

Analysis of Internal Conflict · Universidad del Rosario · Mathew H. Charles, PhD