Navigation – Plan du site

AccueilNuméros84Three Stages in The Acquisition o...

Three Stages in The Acquisition of English Response Tokens: A Window into the Development of Common Ground

Trois étapes dans l’acquisition des énoncés de réponse : une fenêtre sur le développement du common ground
Johannes M. Heim, Júlia Rovira Marí et Martina E. Wiltschko

Résumés

Cet article présente une étude longitudinale des énoncés de réponse d’un enfant anglophone britannique entre les âges de 10 et 60 mois. Il est démontré que les énoncés de réponse servent de fenêtre sur le développement de la notion des savoirs partagés (common ground). Sur la base d’une analyse qualitative et quantitative des énoncés de réponse de l’enfant incluant les marqueurs d’adjacence, les répétitions, les marqueurs d’hésitation, les demandes de réparation, et les éléments verbaux et non verbaux servant de signaux d’écoute et de rétroaction, trois étapes du développement de la notion des savoirs partagés sont identifiées. Initialement, l’enfant marque les énoncés de réponse sans les situer dans son ensemble de croyances (sa posture épistémique). À partir de 28 mois, l’enfant signale son accord en plus de marquer les énoncés de réponse. Ceci est interprété comme signalant l’émergence d’un concept préliminaire des savoirs partagés sans marquer de distinction entre les croyances du locuteur et celles de l’interlocuteur. À partir de 41 mois, l’enfant commence à montrer des formes nuancées d’accord et de désaccord, ce qui reflète une séparation des croyances du locuteur et de celles de l’interlocuteur. Ce n’est qu’à ce stade que le marquage des énoncés de réponse reflète une gestion mature des savoirs partagés et de la conversation.

Haut de page

Texte intégral

Introduction

1The goal of this paper is to explore the question of how children develop an understanding of common ground. We approach this question from a linguistic point of view, which provides us with a unique window into this question. That is, adult language plays an important role in the construction of common ground and conversely, construction and management of common ground is one of the core functions of language. To be sure, common ground is not only constructed through language but is also built via common background and experience. This is certainly also true in young children where common perception and especially joint attention might serve as a precursor for common ground. What language allows us to do, however, is to exchange information and hence to go beyond constructing common ground based on what is available in the immediate perceptual field, i.e., we can share experiences that are not tied to the here and now.

2There are several aspects of language that reflect its role in the construction as well as the management of common ground. For example, in many languages the use of definite and indefinite determiners is sensitive to the distinction between old and new information (I. Heim, 2011; Shukla et al., 2022). As illustrated in [1], at first mention, the article about climate change is introduced by the indefinite article an. This signals that the referent is not yet known to the interlocutor but that it is now introduced in the common ground. In the follow up sentence the same referent is referred to with a definite article which indicates that the referent is old information (i.e., already established in the common ground).

[1]
Several years ago, a group of scientists published
an article about climate change. It didn't take long for many of the predictions made in the article to come true.

3In this article, we explore the question as to what linguistic development may tell us about the development of common ground. There is a caveat, however. Units of language, like definite and indefinite articles, are acquired relatively late. It is well known that children incorrectly use definite articles when referring to a new discourse referent (Karmiloff-Smith, 1979). Nevertheless, an awareness of common ground is present early on. That is, joint attention — a pre-condition for common ground — is found in children long before they become verbal (Carpenter & Tomasello, 1995; H. Clark, 1996; Cosper & Pika, 2024). In fact, there is evidence that children may have at least a rudimentary understanding of common ground as early as 10m (months). At this age, infants repeat the actions of others they observe thereby placing them in the common ground. According to Clark & Bernicot (2008), this repetition of action is a direct predecessor of verbal repetition of new information, as in [2].

[2]
Mother: Where’s the baby duck?
Child (1;5): Duck.
Mother: The baby duck.
(Clark, 2015: 333 (14), cited from Masur & Olson, 2008)

4Clark & Bernicot interpret the child’s repetition as a mechanism for grounding as it allows the child to signal that whatever is repeated is now part of the joint attention between the interlocutors and thus part of the common ground. Hence, linguistic behaviour may give us clues about the development of common ground from the very start. Note crucially that in [2], the child does not use a definite article. In these early stages of acquisition, the child uses nouns and verbs without the scaffolding of the functional elements like articles that later serve as clear indicators of common ground management. Arguably, the acquisition of definiteness appears to lag the development of common ground due to other factors pertaining to the trajectory of language acquisition. Hence, the development of grammar in the classic sense does not appear to be an adequate tool to allow us to draw conclusions about the development of common ground.

5However, linguistic capacities do not only include the kinds of words and morphemes traditionally included in the analysis of sentence grammar and used to construct propositional content (henceforth propositional language). Rather it also includes units of language that are dedicated to the regulation of linguistic interaction, and which include elements that are not necessarily classified as « words ». It includes units of language that are variously referred to as interjections (wow, oh), discourse markers (so, well), response tokens (yeah, mhm), hesitation markers (uh, umm), but also elements that are considered to be words but which can be used in purely interactional ways, such as names used as vocatives (Hey Ismail). Following Wiltschko (2021), we refer to units of language that belong to this domain of language as interactional language.

6What is crucial for our purposes is that interactional language plays an important role in grounding, i.e., the moment-by-moment exchanges that establish information as common ground within a conversation (Clark & Schaefer, 1987, 1989; Clark & Brennan, 1991). And significantly, interactional language is available at the very onset of verbalizations. That is, from the very start, the child not only uses nouns and verbs but also interactional units of language such as requests for responses (huh) as well as response tokens (yeah, no). Given that interactional language is available from the beginning of verbal development, it provides a unique window into the development of common ground independently of the cognitive complexities that come with the development of (in)definite articles, for example.

7The main goal of this paper is to report on a corpus study of the development of a particular subset of interactional language, namely response tokens. Building on recent work on that proposes that linguistic development is best characterized in stages (Perszyk & Waxman, 2018; Friedman et al., 2021), we show that interactional language is available early and gradually gains in complexity in use (see also Shirai et al., 2000; Paul & Yan, 2022; Bosch, 2023). We focus on the response tokens in a child aged 10 months to 5 years and identify the same three stages we have found in the development of other interactional language, such as invariant tags (Heim, J. & Wiltschko, 2025) and variant tags (Heim, J., 2023) in English. At a first stage, response tokens are used merely to signal responding (and thus that the child is participating in verbal interaction). At a second stage, the child increasingly uses response tokens not only to signal responding but also to signal that there is a common ground with the interlocutor. Finally, at a third stage there is an additional dimension available: the child can now also indicate disagreement, which in turn implies that the child can differentiate between her own beliefs and those of her interlocutor.

8The paper is organized as follows. In section 1, we briefly introduce the role of interactional language in adult language with a special focus on response tokens. In addition, we review the findings of J. Heim & Wiltschko (2025) on the development of huh, which serves as an initiation token, and whose development aligns with the three-stage development we propose here. In section 2, we provide the methodological details of our data selection and annotation. In section 3 we provide quantitative and qualitative support for the three stages described above. In section 4, we develop these findings into an account of the development of common ground. We conclude in section 5.

1. Background

9When we talk, we not only use language to exchange information; we also use language to manage this exchange. This is the main purpose of interactional language. Specifically, in adult language, interactional language serves two main functions: it can be used to regulate turn-taking as well as grounding. In this section, we provide a brief overview of interactional language in adult language, with specific emphasis on response tokens (section 1.1). In addition, we summarize the main findings of J. Heim & Wiltschko (2025) on the acquisition of utterance final huh (section 1.2). This will serve as the backdrop for our exploration of the development of response tokens in the present study.

1.1. The role of response tokens for grounding in adults

10In conversational interactions, interlocutors take turns. One of the core insights of conversation analysis (Sacks et al., 1974) is that turn-taking is surprisingly systematic and points to the existence of a communicative competence. At the core of this system are adjacency pairs which can roughly be characterized as consisting of moves that initiate followed by reaction moves (see Wiltschko, 2021 for an overview). For example, the question in [3] is an initiation move, while the subsequent answer constitutes a reaction move.

[3]
Initiation: What time does the movie start?
Reaction: It starts at 9 o’clock.

11Interactional language is typically found at the edges of utterances (i.e., in sentence-peripheral position). Roughly, units of interactional language that are found at the end of an utterance (no in [4])) mark initiation moves while those that are found at the beginning (yeah in [4]) define reaction moves. Specifically, the utterance final no in [4] is a request for confirmation of the propositional content denoted in the host clause whereas the utterance initial yeah indicates agreement with the propositional content of the prior turn.

[4]
Initiation: The movie starts at 8,
no?
Reaction:
Yeah, I think so, too.

12Crucially, in adult language, both confirmationals (markers of initiation) and response tokens, (markers of reaction) can simultaneously serve various functions. Specifically, the use of a confirmational signals not only a request for response but also contributes to marking aspects of the epistemic state of the interlocutor. That is, in [4] the use of the confirmational indicates that the speaker is biased towards believing the propositional content (hence they are not simply asking a question) but they are not certain either (hence they are not simply uttering a declarative clause). Thus, initiation markers simultaneously contribute to turn-taking and grounding.

13Similarly, response tokens are not only used to mark a reaction, they can also contribute to indicating aspects of the speaker's epistemic state. While in [4] yeah indicates agreement with the propositional content of the prior utterance, it is also compatible with disagreement. Specifically, it may be used to simply acknowledge the prior speaker´s utterance, while disagreeing with the propositional content of the prior utterance, as in [5]. This is evident from the fact that yeah can cooccur with no without creating a contradiction (Burridge & Florey, 2002).

[5]
Initiation: That was the best movie ever!
Reaction:
Yeah no, I really don´t think so.

14What will be crucial for our purpose is that response tokens may be found in reaction moves (i.e., when the reacting interlocutor takes a turn which is elicited by the initiator) but they may also be used unelicited without constituting a turn. The latter are sometimes referred to as backchannels (Yngve, 1997) and are typically short vocalizations (yeah, mhm) or even nonverbal cues (head nods) and are thus non-intrusive, as in [6]. They can, however also consist of repetitions of parts of the prior turn, as in [7].

[6]
B: I’ve listen’ to all the things that chu’ve said
An’ I agree with you so much.
Now, I wanna ask you something. I wrote a letter (pause)
A:
Mh hm,
B: t’the governor.
A:
Mh hm ::,
B: -telling ’im what I thought about i(hh)m!
A:
Sh:::!
B: Will I get an answer d’you think,
A:
Yes
(adapted from Schegloff, 1982: 82 [4])

[7]
A: I got everything taken care of. I got insurance on it too.
B: how much it
A: under my name. eleven hundred a year.
B:
eleven hundred.
A: three hundred dollars down
B: that’s cheap man.
(adapted from Clancy
et al., 1996: 361 [3])

15Backchannels are crucial for successful conversations as they signal attention, understanding, and agreement (Clark, 1996). They allow the turn-holder to be sure that the propositional content they wish to add to the common ground is indeed accepted by the listener. And through different backchannels, the listener may additionally indicate how they relate to the proposition. For example, in [6] mhm may simply indicate acceptance, while shh appears to add an emotive dimension, whereas the repetition in [7] might add an element of surprise. In the absence of backchannels, the turn-holder cannot be sure that their utterance is indeed grounded. Thus, backchannels are crucial for the process of grounding (Liesenfeld & Dingemanse, 2022).

1.2. The development of the initiating token huh

16As shown in J. Heim & Wiltschko (2025), initiation tokens, such as invariant tag questions, undergo three stages of development. Here we summarise this three-stage development by discussing the expanding contexts of use of huh in the Brown corpus (Brown, 1973). For reference, adult huh can be used both as an other-initiated repair strategy (Dingemanse et al., 2013) and as an initiating token that seeks to confirm a belief attributed to the addressee. Crucially, it is not well-formed in cases where the speaker is committed to the truth of the propositional content. The use of huh as an other-initiated repair strategy is exemplified in [8] where huh signals the need for the previous speaker to reclaim the floor and reiterate or paraphrase the original contribution. We call this the responding function of huh.

[8]
A: I see you have a new dog.
B:
Huh?

17The repair request in [8] does not include any propositional content. It is purely interactional in that it calls on the addressee to respond without identifying the response target. This is different from initiating huh in [9] which targets the interlocutor’s belief about a proposition. Here, huh seeks confirmation from the addressee that they believe that the dog is new. We therefore call this the grounding function of huh.

[9] You have a new dog, huh?

18But a comparison in [10] with another type of initiation marker (eh) shows that the use conditions of initiating huh are more nuanced than just confirming a speaker belief. While eh can seek confirmation of a belief to which the speaker can commit, huh cannot (as indicated by the asterisk preceding A's utterance in [10]). The latter can only seek confirmation of a belief they associate with the addressee’s ground, not the speaker’s. In other words, only eh, but not huh, seems to have a speaker-oriented function; and separating that from an addressee-oriented function is essential for understand its use.

[10]
A: *I have a new dog,
huh?
A’: I have a new dog,
eh?
B: Of course, you do! Sorry for not mentioning anything.

19In child language development, the three functions of interactional language — responding, grounding, and separating perspectives — are incrementally acquired in three developmental stages. For the acquisition of huh this translates into the following trajectory:

Stage I:
huh as a request for response
Stage II:
huh as a request for confirmation of a belief held by both interlocutors
Stage III:
huh as a request for confirmation of a belief held by one interlocutor

20These three stages constitute an expansion from the initial response function of huh into interlocutor-specific functions. Before the child has access to dedicated speaker- and addressee-oriented functions, there is an intermediate stage where the child does not sperate their beliefs. Below, we provide examples of huh representative of each stage of development, which in turn supports the proposal of a stage-wise expansion of its function. The examples are from Adam and Sarah, two of the children from the Brown corpus (Brown, 1973).

[11]
Adam: Where go,
huh? (2;07)
Mother: I don’t know.

[12]
Sarah: That look nice,
huh? (3;05)
Ken: Very nice.

[13]
Sarah: We got Grampy socks,
huh? (4:10)
Mother: You bought Grampy socks?
Sarah: Yeah.

21Huh in [11] exhibits a non-adult combination with a wh-question to mark the request for response; the example in [12] includes a subjective layer typical for adult usage. Occurrences at this intermediate stage nevertheless lack the separation of addressee- and speaker-oriented perspective we find in [13]. There, the child asks for a confirmation of her own beliefs about an action of the interlocutor. Interestingly, the speaker-oriented use of huh in [13] is not adult-like as it maps onto the use-conditions of eh, not huh. Adult huh cannot seek confirmation of a belief to which the speaker has committed, but Sara’s response to her mother’s question reveals that this is exactly the use of huh in [13]. It seems, therefore, that even when all interactional functions are available to the child, they still need to learn when to use eh, and when to use huh. In other words, they need to learn to distinguish what is logically possible from what is attested in the target language.

22Based on these observations about the initiation token huh, we expect a similar developmental trajectory for response tokens, including backchannels, interjections, repair requests, and hesitation markers. That is, we take the developmental path of huh to be representative of the acquisition of other interactional language. Initially, the child cannot access the interactional functions required to monitor interlocutor beliefs. Units of interactional language are therefore reappropriated for more basic functions, such as managing turn-taking, which is available from birth (Dominguez et al., 2016). Only with further cognitive maturation can we expect to see an increasingly nuanced management of common ground (see Section 4 for further details). In Section 2, we provide the details of how we went about collecting and annotating our data, which we then present in Section 3.

2. Methods

2.1. Corpus choice

23To investigate the development of response marking in English at a longitudinal scale, we chose the Sekali corpus (Morgenstern et al., 2018) from the British English collection on CHILDES (MacWhinney, 2000). Specifically, we studied a monolingual child, Ellie, who was raised by an English-speaking family from Warwickshire, UK. The corpus is based on monthly one-hour-long audiovisual recordings from the age of 0;10 to 3;06 at which point recordings were reduced to a bimonthly rhythm until 5;0 when Ellie’s sibling was born. Overall, this afforded us with 43 recording sessions with Ellie alone. Most videos were recorded by Ellie’s grandmother, a special needs teacher. Ellie’s mother worked as a research technician in biotechnology at the time of recording, her father as an electronic engineer. The corpus was primarily chosen because of the nature of the available data: Only videotaped conversations provide sufficient contextual information to help annotators disambiguate the use conventions of multifunctional units of interactional language, such as response tokens.

2.2. Data selection

24Because we aimed for a comprehensive investigation of the development of response tokens, we cast our net widely by incorporating various response marking strategies. We thereby follow Sbranna et al. (2024) in distinguishing two types of short responses. i) uninitiated response tokens, which are those that engage with a previous speaker’s utterance without an expectation of a turn change by the original speaker and ii) initiated response tokens, which are those that involve an expected turn takeover after a question, indirect request, declarative question or tag question.

25Response tokens were identified by the second author through watching the recordings alongside the automatically generated transcripts provided on CHILDES. To capture all forms of responding, we expanded traditional forms of responding to include hesitation markers, interjections, partial repetitions, and repair requests. Our selection included 803 interactional units of language, of which 36 had to be excluded because the annotators could not agree on a label (often because the child employed preverbal units of language). Among the remaining 766 items, 330 included a turn change, the remaining 436 did not.

2.3. Annotation procedure

26All 766 items were annotated by the second author during the selection process; secondary annotation was completed by the first author relying on contexts consisting of up to two previous and following transcript lines. In this secondary annotation, original videos were only consulted when the context was ambiguous. The final annotation scheme consisted of the categories listed below:

• verbal backchannels: brief utterances that do not constitute a turn takeover and which don't contain (substantial) propositional content
• non-verbal backchannels: nodding, smiling, laughter, etc.
brief repetitions of elements in the previous speaker’s turn
• adjacency markers: brief, elicited response to a previous speaker
• interjections mark a change in belief or attention (e.g., oh, look)
• repair requests (e.g., huh, eh, or pardon)
• hesitation markers (e.g., uhm, uh, and hmm) in response to a question.

2.4. Inter-annotator agreement

27Annotators agreed on 536 labels among the 803 selected items in their annotations (66.7%). 195 of these 240 disagreements were based on different criteria for identifying verbal backchannels. One annotator focused on the lack of floor-holding; the other on the turn interruption through a backchannel. This had consequences for those response tokens where a third interlocutor would speak after the potential backchannel. Another difference centred around the role of non-canonical tag questions and vocatives, which counted as potentially initiating a turn-change. We decided to exclude any form of initiation for backchannels and allow involvement of a third speaker. A total of 36 items had to be excluded because annotators could not identify an unambiguous use or because the verbalisation was unintelligible. The remaining 70 disagreements were resolved through discussion and reanalysis by the two annotators.

3. Results

28In this section, we present the onset and trajectory of the 766 response tokens that were finally analysed. We begin with an overview of the first occurrences of the response tokens as well as the trajectory of lexical diversification in each of the annotated categories over time. For both initiated and uninitiated response tokens, this development appears to come in three stages. We then provide some examples that serve as evidence for a qualitative difference between these stages whereby the child initially just reacts to the propositional content of a previous utterance and later adds an increasingly nuanced perspective independently of whether the response is initiated or not. We argue that these three states reflect the development of the concept of the common ground in the child’s development. Specifically, the use of the response tokens which the child displays shows a growing ability to construct and manage common ground and, significantly, to separate speaker and address beliefs.

3.1. The onset and trajectory of initiated and uninitiated response tokens

29Ellie uses response tokens from the onset of recordings. Even at 10m, when only a few isolated words occur in her output, she already uses interactional language, such as hi and bye in response to others and backchannels via nodding and pre-word verbalisation. Figure 1 provides an overview of Ellie’s brief responses in the transcripts, ordered by age and lexical item, independently of context of use. Dashed lines mark two notable increases in response marking, one at 28 months, and a second at 41 months.

Figure 1: Count and lexical diversification of response tokens (n>1%) across age in months

Figure 1: Count and lexical diversification of response tokens (n>1%) across age in months

30The increase at 28m is more subtle than the one at 41m. Yet even the former includes a notable increase in the use of yeah, and the onset of the hesitation marker um. This demonstrates that Ellie already employs verbal strategies to structure conversation. While the number of individual response tokens fluctuate after the beginning of the new stage at 28m, response marking is consistently higher than before. Even when we separate the contexts of use of the various response tokens depending on whether or not they are initiated by the previous speaker, the overall pattern remains the same. In both categories, we observe an increase in token frequency and lexical diversity that proceeds via stages.

Figure 2: Count of various response tokens (n>1%) across age in months, separated by context of use

Figure 2: Count of various response tokens (n>1%) across age in months, separated by context of use

31The comparison of initiated and uninitiated response tokens strongly suggests that initiated response tokens lead the way in response development. Initiated tokens are more frequent at the beginning of recordings and diversify more quickly. Nevertheless, both types show a change in behaviour around the 28m and 41m mark. For instance, yeah and no are not used in either initiated or uninitiated responses until 29m. Similarly, hesitation markers, like um, which signal the intention to maintain the turn, do not occur until Stage II.

32A more comprehensive overview of the different response strategies is provided in Table 1, which shows that all types of responses are available within the first 18m, except for hesitation marking. The incremental growth by stage is particularly evident in adjacency pairs and verbal backchannels, but hesitation markers and interjections also increase notably.

Table 1: First uses and frequency per stage of response tokens by type

 

Adjacency pairs

Non-verbal back-channel

Interjection

Repetition

Verbal back-channel

Repair

Hesitation marker

Total

1st use

10m

10m

12m

15m

16m

17m

29m

Stage I

27

7

8

36

10

2

90

Stage II

52

7

6

24

58

3

11

161

Stage III

168

4

26

35

257

2

23

515

Total

247

18

40

95

325

7

34

766

33Further support for an incremental growth pattern comes from analysing the proportion of response strategies across stages (Figure 3). In line with the overall growth in vocabulary size, Ellie can draw on more and more strategies to engage with her interlocutors. And as the proportion of verbal backchannels increases, the proportion of non-verbal backchannels decreases, while adjacency pairs remain stable throughout. The early presence of interjections and non-verbal backchanneling shows, however, that Ellie engages with her interlocutors from early on in her language development, even when a turn-take is not expected.

Figure 3: Proportion of response types per stage

Figure 3: Proportion of response types per stage

34The increasing frequency of repetitions, although perfectly in line with the overall pattern, requires further scrutiny due to their multifunctionality. Specifically, repetitions can serve to (i) repeat and rehearse the label for a new word, (ii) respond to a question by repeating the relevant information to confirm an alternative, or (iii) backchannel by repeating part of the previous turn. Ellie pursues all three strategies, albeit at different proportions at different stages. Among her responses, we found 93 repetitions. Of these, 28 are repetitions of some words of the previous utterance without identifiable pragmatic purpose; 23 serve as responses to a question; and 39 are employed to backchannel. Note that the use of repetitions changes over time. Figure 4 shows that these uses broadly change in lockstep with the development of response marking as described above. There is little backchanneling via repetitions during Stage I (i.e., before 28m) and a considerable increase of that use at Stage III (from 41m).

Figure 4: Proportion of repetition use per six-month window by use

Figure 4: Proportion of repetition use per six-month window by use

35The fact that the use of repetitions largely maps onto the three stages discussed above supports the proposal that response marking develops in stages. In the case of repetitions, it is the change in use, not the lexical variation, that aligns with the stages.

3.2. Three stages of marking responses

3.2.1. Stage I – Marking response (10-27m)

36Responding in conversation requires constant monitoring of the interlocutor's turn. This is because turn-constructional units must be identified so as to react with little delay or overlap (Sacks et al., 1974). The complexity of this task makes it somewhat surprising to find Ellie produce instances of verbal backchanneling as early as 10m. At a point where she can only say a few words, Ellie knows if and when backchanneling is appropriate. What stands out at this early stage is both the paucity of variation in backchannels and the frequent presence of non-verbal backchanneling, such as nods and laughter, as well as vocalisations that resemble repetitions. Before the end of stage I, Ellie only uses interjections like oh (from 16m; see Table 2) or uh oh (from 20m), and positive response particles like yeah (from 20m) and yes (from 21m). These continue to be her dominant forms of backchanneling until significantly later. Negative response tokens are exclusively used in initiated responses to questions, while positive ones are available for both initiated and uninitiated responding.

Table 2: First occurrences of elicited and unelicited response tokens

10m

11m

12m

14m

16m

19m

20m

21m

27m

initiated

bye

hi

no

thank you

oh

huh

yeah

yes

8

uninitiated

wow

uh oh

yeah

yes

okay

5

2

1

1

1

1

1

1

2

1

13

37Example [14] exhibits a response to a previous turn without initiation. The time gap between turns is (almost) adult-like (here: 410ms; with an adult average of around 200ms (Stievers et al., 2009). It is certainly short enough to warrant the assumption that Ellie does not wait until the end of her grandmother’s turn to launch the articulation of her response (estimated at 600ms, Levinson & Torreira, 2015).

[14]
Grandmother: We need to open the dishwasher.
Ellie (1;08): Yeah.
Grandmother: Wait one minute though.

38Thus, even simple reactions such as in [14] demonstrate that Ellie has complex interactional skills at a point where multiword utterances (the accepted marker for complex representations) are still rare.

  • 1 We follow the CHILDES convention of marking turn overlaps with squared bracketing.

39Next consider example [15]. Here Ellie combines verbal and non-verbal backchanneling while the mother continues to speak. The timing coincides with a syntactic and prosodic boundary. Again, this supports the view that Ellie recognises turn constructional units to time her backchannelling appropriately (i.e., right where the mother signals continuation).1

[15]
Mother:
Should we go outside…
[Ellie (1;11): [nods] “Yes.”]
…and come when you go out.

40In brief, Stage I is characterised by brief reactions at the expected points in conversation, sometimes supplemented or substituted by non- or pre-verbal response marking in line with the overall linguistic development of the child. At no point during this stage, however, do we find evidence of a monitoring or attempts of expanding the common ground.

3.2.2. Stage II – Marking agreement (28-40m)

From 28m, we find a notable increase in Ellie’s use of both initiated and uninitiated response tokens (see Figures 1-3). She exhibits more diversity in lexical choices, including multiword units such as you think (31m) and I know (38m). Table 3 shows that this increase in lexical diversity is particularly present in initiated response tokens, but even the few additions of uninitiated tokens afford the child with the possibility to show agreement. Hence, we observe the beginning of response tokens reflecting the interlocutor's beliefs. Due to their subjective nature, we take these to be first markers of common ground management. The arrival of hesitation markers (28m), which are used to signal that the speaker is planning their next utterance (Clark & Fox Tree, 2002), further shows that children possess a growing metalinguistic awareness. Nonverbal backchannels continue to be present at this stage.

Table 3: First occurrences of various units of initiated and uninitiated response tokens

28m

29m

31m

32m

33m

35m

36m

38m

39m

40m

initiated

no

um

mm

you think

uh

oh yeah

mhm

hmm

eh

hey

right

aye

I know

14

uninitiated

um

mhm

hm

uh

I know

well

huh

eh

7

2

2

7

2

3

1

1

2

2

1

21

41To see this development, consider some examples. In [16], Ellie’s use of you think shows that her response goes beyond a simple reaction to propositional content. The mother’s preceding utterance calls for Ellie’s attention, but Ellie’s response includes subjective judgments. The mother’s next turn shows, however, that Ellie’s backchannel is not interpreted as a questioning of the mother’s original belief.

[16]
Mother: I’m galloping fast Ellie, look. Whoa!
Ellie (2;07):
You think?
Mother: Can you go that fast.

42These early traces of perspective taking extend to agreement with subjective judgments even when the mother does not call for a confirmation. Ellie seems to have learned that backchanneling can make the other person feel understood, which further demonstrates that she understands the importance of building common ground. This holds independently of whether the response is initiated or uninitiated, as shown in [17].

[17]
Mother: She’s a bit silly sometimes.

Ellie (3; 03):
Yeah.
Mother: When she wanted to play with you Sylvanian families, didn’t she?
Ellie:
Yeah.

43Repetition, tag questions, and hesitation markers also contribute to managing dialogue. The increase in quantity of response marking from 28m therefore is accompanied by an increase in lexical diversity and one in quality: besides simply responding, Ellie is now also able to signal shared beliefs and actively manages common ground.

3.2.3. Stage III – Separate perspectives (from month 41)

44From 41m, Ellie has a notable increase of verbal backchanneling. We see the arrival of 10 further response tokens in uninitiated contexts, which allow her to backchannel with greater nuance. The fact that uninitiated response tokens increase to a larger degree affirms the hypothesis that initiated responding, which displayed a similar expansion earlier, leads the development of response marking. Table 4 lists the arrival of all added response tokens across both contexts of use. Significantly, however, all response tokens strongly increase in frequency.

Table 4: First occurrences of elicited and unelicited response tokens

41m

44m

46m

48m

50m

53m

57m

60m

initiated

hm

uhm

cheers

pardon

4

uninitiated

look

uhm

yay

good

oops

mm

ha

ah

tsk

hooray

yum

11

5

1

2

2

1

1

1

2

15

45Moreover, we note that yeah is by far the most frequent response token (n = 305) across Ellie’s transcripts, but it occurs particularly often at 41m (n = 48), which corresponds to the onset of Stage III (see Figure 1). A possible explanation for the outstanding increase of yeah at 41m is that Ellie unlocks a different understanding of the use of backchannels, which she rehearses during that month. The phenomenon of rehearsing new capacities is familiar from other areas of language development (Weir, 1962). Example [18] shows how wide-ranging the use of yeah can be now: it can signal agreement to a fact, as in the first two uses, and to a subjective judgment that relates to the experience of the child (the waves were particularly big for 3-year-old Ellie). It is also worth noting that the grandmother to whom Ellie is responding seems to have been an observer only to the original experience.

[18]
Grandmother: You did! With mommy
[Ellie (3;05):
yeah]
holding on tight.
Ellie:
Yeah!
Grandmother: Cause they were rather big those waves.
Ellie:
Yeah!
Grandmother: Pushing you above them.

46At this stage, there are also instances of initiated responses that show a good command of how to employ positive and negative response markers as in [19], again at a subjective level.

[19]
Friend: The troll doesn't look very nice either, does he?

Ellie (4;00):
No.
Friend: But you made a really good story for him, didn’t you?
Ellie:
Yeah.

47At this stage, Ellie seems to be able to clearly distinguish speaker from hearer perspective in her response marking. Consider [20] where she first confirms the perception of the original speaker and then confirms it again by switching perspective to her own perception.

[20]
Grandmother: Ellie was asleep.
Ellie (4;05):
Yeah.
Grandmother: Mhm?
Ellie: And I was tired.

48While verbal backchannels increase in frequency together with interjections, non-verbal backchannels and repetitions continue to be used to signal engagement without claiming the floor. Examples [21] and [22] point to the wide range of response tokens the child has now available. The repetition in [21] shows that reaction to a previous turn is available just as is redirecting the interlocutor’s attention and confirming a belief, as in [22].

[21]
Grandmother: Hold it in your hands.
Grandmother: And scrunch it up a bit.
Ellie (3;06):
Scrunch!
Grandmother: Squid it.
Grandmother: Go on.

[22]
Mother: You need a two Ellie…
[Ellie (4;09):
tsk]
…but look you've got three takeaway away one.
Mother: Like I had four takeaway one.
Ellie:
Oh, look.
Mother: I know but you need a two, don’t you.
Ellie:
Yeah.
Mother: Ah, not my three!

49In sum, we observe that, at Stage III, Ellie is able to use a large variety of response tokens. They range from merely signaling a reaction to a previous utterance to agreeing and disagreeing with subjective perspectives. In brief, Ellie now has the abilities to use response tokens with all the functions available in adult language.

4. Discussion

4.1. Three stages in the development of grounding

50Our findings show that managing the common ground via response markers develops in three stages: i) affirming reactions; ii) signaling shared agreement; and iii) nuanced perspective taking on a previous speaker’s turn.

51Evidence for a three-stage development comes from the increases in frequency of response tokens at 28m and 41m, the lexical diversification of initiated response tokens at Stage II and uninitiated response tokens at Stage III, as well as the proportional change in the context of use of repetitions that mirror the development of dedicated response markers.

  • 2 We assume that first instances of protests (including through the use of no) do not constitute a co (...)

52What does this mean for the development of common ground? We propose that each stage in the development of response marking corresponds to an expansion of the ability to participate in constructing and managing common ground. Figure 4 visualises the growing response options to reflect this development. At Stage I, the child does not yet participate in constructing the common ground. The first units of interactional language mark the most essential strategies of linguistic interaction, i.e., turn-taking. This consists of signaling a response (e.g., yeah) as we showed here as well as requesting a response (via utterance final huh), as observed in J. Heim & Wiltschko (2025). From Stage II onwards, however, responding goes beyond merely signaling a reaction to a previous turn. The child can now express propositional attitudes, albeit with a default assumption that these attitudes are shared between speaker and addressee.2 In other words, there is a generalized common ground. The assumption correlates with the absence of a fully developed Theory of Mind, which is required to separate interlocutor perspectives and represent them as possibly diverging worlds (Wimmer & Perner, 1983; DeVilliers, 2021).

Figure 5: Three stages toward grounding

Figure 5: Three stages toward grounding

53From Stage III onwards, the child has all three options available: reacting, agreeing, and signaling divergent beliefs. Our data on the expanding contexts of use of repetitions lends further support to this analysis: while language learning (through repeating new labels) dominates Stage I and decreases over time, confirming shared beliefs dominates Stage II, which in turn is superseded by using repetitions primarily for backchanneling from Stage III, just as adult speakers do. This trajectory does not mean that the previous functions stop being available, however. The more likely explanation may be that every added function emerges out of the previous ones due to a greater diversification of its contexts of use, in line with Clark & Bernicot’s (2008) assumption that repetition often is a precursor of things to come.

4.2. Are we there yet?

54Despite the nuanced response marking present at Stage III, Ellie still uses response tokens that are not adult-like in a particular context. Example [23] includes a response token added at Stage III which seeks agreement of a belief. The initiated okay would only be felicitous for an adult speaker in this context if they want to withhold agreement and just acknowledge the speech act. Similarly, example [24] includes an uninitiated response token that may be felicitous for an emotionally disengaged speaker, but not in the present context where the interlocutors agree.

[23]
Grandmother: It’s big, isn’t it?

Elle (3;08):
Okay.
Mother: For children, I think that can go.

[24]
Grandmother: I've forgotten what I've got in here
.
Grandmother: Incredible!
Ellie (3;08):
Good.
Grandmother: Such a long time since you all played with it.

55We submit that children continue to refine the use of response marking after they have learned all logically available options — acknowledging, (dis)agreeing with a common belief, and signaling a possible divergence of beliefs. We have independent evidence for such ‘pruning’ of possible, but not attested uses from other interactional units of language, such as huh. Specifically, in child language huh continues to exhibit a context of use not available in the adult use of huh. Rather the child may use huh in the same contexts where adults must use other units of language (e.g., Canadian eh [J. Heim & Wiltschko, 2025]). The pruning process will then gradually lead to a full alignment of what could be used in a particular context given the available options, and what is attested in the child’s target language.

4.3. Lessons for propositional language

56We began by pointing to the limitations of studying the development of common ground through the lens of propositional language due to the late mapping of definite and indefinite articles onto given/new distinctions. We then argued for a three-stage development of common ground based on the observed three stages of the acquisition of interactional language, specifically the acquisition of initiated and uninitiated response tokens. We now return to propositional language and briefly spell out the predictions we can make for its relation to common ground in language development based on our observations about response marking.

Figure 6: Parallel development of interactional and propositional language

Figure 6: Parallel development of interactional and propositional language

57Based on the alignment of developmental stages and functions of interactional language represented in Figure 5, we predict that propositional language also expands from basic information exchange to a nuanced anchoring of this information into the world as represented in Figure 6. Correspondingly, early propositional language prioritises labelling and exchanging information about the world, just as we have seen in the uses of repetitions at Stage I. Here, engaging with the interlocutor serves primarily to partake in basic communication without negotiating beliefs or attitudes. At Stage II, we expect the child to loosely anchor propositional language in the world at large. Just as shared beliefs begin to emerge, and repetitions predominantly serve to confirm beliefs, so does the child begin to evaluate whether propositional content is true or false. It is at this stage that we see the arrival of inflecting for tense and aspect, albeit with frequent omissions and commissions.

58The concept of underspecified categories may also help address the conundrum about explaining the late acquisition of definite and indefinite articles mentioned at the outset of this paper and their relation to common ground development. Schaeffer & Matthewson (2005) show that child English frequently exhibits definite articles in contexts where the addressee is unfamiliar with a referent. In adult English, an indefinite article is obligatory in such contexts. Schaeffer & Matthewson (2005) propose a concept of non-shared assumptions to explain the overgeneralisation observed in children. In contrast, we propose that children fail to understand the link between givenness and definiteness because common ground is not yet fully developed. At this stage, children fail to incorporate different perspectives and therefore assume that what is known to them is also known to the addressee. This affects the expression of (in)definite determiners. Finally, at Stage III, propositional language develops fully and thus allows for reported events to be anchored to the utterance situation (in English expressed via tense morphology) as well as to be presented from a particular point of view (in English expressed through aspectual). Nevertheless, there will still have to be a final step of separating what is logically possible from what is attested in the adult language. An obvious context where this process of pruning is evidenced in propositional language is the realisation of past tense morphemes in English: before a child’s output matches the adult language, the child needs to understand how past tense is realised. At this point, children apply the default realisation too broadly, until they finally can distinguish between regular and irregular forms as in the adult language (see e.g., Marcus et al., 1992).

Conclusion

59In this paper, we argued for three stages in the development of common ground based on the emergence of different response tokens in the output of an English-speaking child aged 10m to 60m. Across initiated and uninitiated contexts we observed developmental trajectory comparable to the one observed in the context of initiating tokens, such as variant and invariant tags (J. Heim, 2023, J. Heim & Wiltschko, 2025). Specifically, we found that Ellie from the Sekali corpus (Morgenstern et al., 2018) initially uses non-verbal backchannels, repetitions, and interjections to simply mark responses in uninitiated contexts alongside a few positive response markers following questions, vocatives and requests. These early response tokens are used to react to a previous turn without incorporating a subjective attitude. This changes around 28m with the arrival of new response tokens, particularly in initiated contexts. At this second stage, Ellie now also provides a perspective on propositional content that grounds in the assumption that speaker and addressee share the beliefs held by the child. We conceive of this as a precursor to the notion of common ground. However, at this stage it appears that speaker and addressee beliefs are not yet separated, and agreement is the default. Stage III has its onset at 41m with a notable burst in response marking and a lexical diversification of responding particularly in uninitiated contexts. With a large inventory of response tokens, the child is able to provide a nuanced perspective on propositional content that can also communicate disagreement. Because speaker and addressee beliefs can now be conceived of as separate entities, the child has arrived at an adult-like concept of common ground. The only remaining task for the child is to check logically possible ways of responding against what it attested in the target language. She has to prune lexical connections that overextend the possible contexts of use for individual response tokens. By investigating the development of common ground through the lens of response tokens, we were able to work around the restrictions of propositional strategies of common ground construction and management due to their late arrival in language development.

Haut de page

Bibliographie

Brown R., 1973, “Development of the first language in the human species”, American psychologist 28(2), 97–106. DOI: https://doi.org/10.1037/h0034209

Bosch N., 2023, “Not all complementisers are late: a first look at the acquisition of illocutionary complementisers in Catalan and Spanish”, Isogloss 9(1), 1-39. DOI: https://doi.org/10.5565/rev/isogloss.313

Burridge K., & Florey M., 2002, “‘Yeah-no he’s a good kid’: A discourse analysis of yeah-no in Australian English”, Australian Journal of Linguistics 22(2), 149–171.

Carpenter, M., & Tomasello, M., 1995, “Joint attention and imitative learning in children, chimpanzees, and enculturated chimpanzees”, Social Development 4(3), 217-237.

Clancy P. M., Thompson S. A., Suzuki R., & Tao H., 1996, “The conversational use of reactive tokens in English, Japanese, and Mandarin”, Journal of pragmatics 26(3), 355-387.

Clark E. V., 2015, “Common ground”, in B. MacWhinney, & W. O’Grady (eds), The handbook of language emergence, 328-353.

Clark E. V., & Bernicot J., 2008, “Repetition as ratification: How parents and children place information in common ground”, Journal of child language 35(2), 349-371.

Clark H. H. 1996, Using Language, Cambridge, Cambridge University Press.

Clark H. H., & Brennan S. E., 1991, “Grounding in communication”, in L. B. Resnick, J. M. Levine, & S. D. Teasley (eds), Perspectives on socially shared cognition, American Psychological Association, 127–149. DOI: https://doi.org/10.1037/10096-006

Clark, H. H. & Fox-Tree J. E., 2002, “Using uh and um in spontaneous speaking”, Cognition 84(1), 73-111.

Clark H. H., & Schaefer E. F, 1989, “Contributing to discourse”, Cognitive science 13(2), 259-294.

Clark H. H., & Schaefer E. F., 1987, “Collaborating on contributions to conversations”, Language and cognitive processes 2(1), 19-41.

Cosper S. H., & Pika S., 2024, “Human turn-taking development: A multi-faceted review of turn-taking comprehension and production in the first years of life”, accessed 12-04-2025. DOI: https://doi.org/10.31234/osf.io/yjad2

deVilliers J. G, 2021, “The role (s) of language in theory of mind”, in M. Gilead, & K. N. Ochsner (eds), The neural basis of mentalizing, Cham, Springer International Publishing, 423-448.

Dimroth C., 2010, “The Acquisition of Negation”, in L. R. Horn (ed.), The Expression of Negation, Berlin, New York, De Gruyter Mouton, 39-72. DOI: https://doi.org/10.1515/9783110219302.39

Dingemanse M., Torreira F., & Enfield N. J., 2013, “Is ‘Huh?’ a universal word? Conversational infrastructure and the convergent evolution of linguistic items”, PloS one 8(11). DOI: 0078273

Dominguez, S., Devouche, E., Apter, G., & Gratier, M., 2016, “The roots of turn‐taking in the neonatal period”, Infant and Child development 25(3), 240-255. DOI: 10.1002/icd.1976

Friedmann, N., Belletti, A., & Rizzi, L., 2021, “Growing trees: The acquisition of the left periphery”, Glossa 6(1). DOI: https://doi.org/10.16995/glossa.5877

Heim Irene., 2011, “Definiteness and indefiniteness”, in K. von Heusinger, C. Maienborn, & P. Portner (eds), Semantics: An international handbook of natural language meaning, 996–1025. Berlin, Boston, De Gruyter. DOI: 10.1515/9783110589443-002

Heim J. M., & Wiltschko M. E, 2025, “Rethinking structural growth: Insights from the acquisition of interactional language”, Glossa 10(1). DOI: 10.16995/glossa.16396

Heim J. M., 2023, “Acquisitional trajectories of North American vs British English tags”, Paper present at LAGB 2023, accessed 12-04-2025. URL: https://hdl.handle.net/2164/22533

Karmiloff-Smith A., 1979, A Functional Approach to Child Language: A Study of Determiners and Reference, Cambridge, Cambridge University Press.

Levinson S. C., & Torreira F., 2015, “Timing in turn-taking and its implications for processing models of language”, Frontiers in psychology 6, 731.

Liesenfeld A., & Dingemanse M., 2022, “Bottom-up discovery of structure and variation in response tokens (‘backchannels’) across diverse languages”, Interspeech 2022, 1126-1130.

MacWhinney B., 2000, The CHILDES project: The database vol. 2, Psychology Press.

Marcus G. F., Pinker S., Ullman M., Hollander M., Rosen T. J., & Xu F., 1992, Overregularization in Language Acquisition. Monographs of the Society for Research in Child Development 228, 57(4), Hoboken, NJ, Wiley.

Masur E. F., & Olson J., 2008, “Mothers’ and infants’ responses to their partners’ spontaneous action and vocal/verbal imitation”, Infant behavior and Development 31(4), 704-715.

Morgenstern A., Blondel M., Beaupoil-Hourdel P., Benazzo S., Boutet D., Kochan A., & Limousin F., 2018, “The blossoming of negation in gesture, sign and oral productions”, in M. Hickman, E. Veneziano, & H. Jisa (eds), Sources of Variation in First Language Acquisition: Languages, contexts, and learners, 339–364.

Paul W., & Yan S., 2022, “Sentence-final particles in Mandarin Chinese”, Discourse Particles: Syntactic, Semantic, Pragmatic, and Historical Aspects 276, 179. DOI: http://doi.org/10.1075/la.276.07pau.

Perszyk, D. R., & Waxman S. R., 2018, “Linking language and cognition in infancy”, Annual review of psychology 69(1), 231-250.

Sacks H., Schegloff E. A., & Jefferson G., 1974, “A simplest systematics for the organization of turn-taking for conversation”, Language 50(4), 696-735. DOI: 10.2307/412243

Sbranna S., Wehrle S., & Grice M., 2024, “A multi-dimensional analysis of backchannels in L1 German, L1 Italian and L2 German”, Language, Interaction and Acquisition 15(2), 243-277.

Schaeffer J., & Matthewson L., 2005, “Grammar and pragmatics in the acquisition of article systems”, Natural Language & Linguistic Theory 23(1), 53-101.

Schegloff E. A., 1982, “Discourse as an interactional achievement: Some uses of ‘uh huh’ and other things that come between sentences”, in D. Tannen (ed.), Analyzing discourse: text and talk. Georgetown University Roundtable on Languages and Linguistics, Washington, DC, Georgetown University Press, 71-93.

Shirai J. & Shirai H., & Furuta Y, 2000, “Acquisition of Sentence-Final Particles in Japanese”, in M. Perkins, & S. Howard (eds), New Directions on Language Development And Disorders, Boston, MA, Springer US, 243–250. DOI: 10.1007/978-1-4615-4157-8_23

Shukla V., Long M., & Rubio-Fernandez P., 2022, “Children’s acquisition of new/given markers in English, Hindi, Mandinka and Spanish: Exploring the effect of optionality during grammaticalization”, Glossa Psycholinguistics 1(1). DOI: 10.5070/G6011120

Stern C., & Stern W., 1928 [1975], Die Kindersprache. Eine psychologische und sprachtheoretische Untersuchung, Darmstadt, Wissenschaftliche Buchgesellschaft.

Stivers T., Enfield N. J., Brown P., Englert C., Hayashi M., Heinemann T., … & Levinson S. C., 2009, “Universals and cultural variation in turn-taking in conversation”, Proceedings of the National Academy of Sciences 106(26), 10587-10592.

Weir R., 1962, Language in the crib, The Hague, Mouton.

Wiltschko M., 2021, The grammar of interactional language, Cambridge University Press.

Wimmer H., & Perner J., 1983, “Beliefs about beliefs: Representation and constraining function of wrong beliefs in young children's understanding of deception”, Cognition 13(1), 103-128.

Yngve V., 1997, “On getting a word in edgewise”, in Papers from the Sixth Regional Meeting, Chicago Linguistic Society, 567–578.

Haut de page

Notes

1 We follow the CHILDES convention of marking turn overlaps with squared bracketing.

2 We assume that first instances of protests (including through the use of no) do not constitute a counter-example. Rather, following Dimroth (2010), the initial use of negation is a matter of refusal rather than a disagreement with a belief. This insight goes back to Stern & Stern (1928: 267) who suggest that “Children’s first no does not mean ‘no, this is not the case’ and thus does not constitute a negative judgment, but means ‘no, it should not be like this’, ‘no, I don’t want that’, ‘no, you should not do this’, and therefore presents a rejecting statement.” (translated from German by and cited from Dimroth, 2010, 42)

Haut de page

Table des illustrations

Titre Figure 1: Count and lexical diversification of response tokens (n>1%) across age in months
URL http://journals.openedition.org/praxematique/docannexe/image/10146/img-1.jpg
Fichier image/jpeg, 62k
Titre Figure 2: Count of various response tokens (n>1%) across age in months, separated by context of use
URL http://journals.openedition.org/praxematique/docannexe/image/10146/img-2.jpg
Fichier image/jpeg, 65k
Titre Figure 3: Proportion of response types per stage
URL http://journals.openedition.org/praxematique/docannexe/image/10146/img-3.jpg
Fichier image/jpeg, 38k
Titre Figure 4: Proportion of repetition use per six-month window by use
URL http://journals.openedition.org/praxematique/docannexe/image/10146/img-4.jpg
Fichier image/jpeg, 32k
Titre Figure 5: Three stages toward grounding
URL http://journals.openedition.org/praxematique/docannexe/image/10146/img-5.jpg
Fichier image/jpeg, 46k
Titre Figure 6: Parallel development of interactional and propositional language
URL http://journals.openedition.org/praxematique/docannexe/image/10146/img-6.jpg
Fichier image/jpeg, 40k
Haut de page

Pour citer cet article

Référence électronique

Johannes M. Heim, Júlia Rovira Marí et Martina E. Wiltschko, « Three Stages in The Acquisition of English Response Tokens: A Window into the Development of Common Ground »Cahiers de praxématique [En ligne], 84 | 2025, mis en ligne le 10 décembre 2025, consulté le 11 août 2026. URL : http://journals.openedition.org/praxematique/10146 ; DOI : https://doi.org/10.4000/15b3u

Haut de page

Droits d’auteur

CC-BY-NC-ND-4.0

Le texte seul est utilisable sous licence CC BY-NC-ND 4.0. Les autres éléments (illustrations, fichiers annexes importés) sont susceptibles d’être soumis à des autorisations d’usage spécifiques.

Haut de page
Rechercher dans OpenEdition Search

Vous allez être redirigé vers OpenEdition Search