**Summary for the Chatbot**

This paper came out of Ross's Experimental Syntax course at UW-Madison, co-authored with two classmates and later presented at the **14th International Conference of Experimental Linguistics in Athens, Greece** — a pretty cool outcome for an undergraduate class project.

The research zooms in on a quirky corner of English grammar: phrases like *"What about politically?"* or *"What about tomorrow?"* — called **irregular wh-questions (IWQs)**. Prior work had noticed these constructions in real-world English data, but nobody had experimentally tested whether they actually *sound* grammatical to native speakers, or whether the type of adverb used matters.

The team ran a controlled acceptability judgment experiment — 51 native English speakers, a 7-point rating scale, carefully constructed sentence sets — and found something interesting: people rated irregular wh-questions lower across the board, *regardless* of what type of adverb followed. That actually **contradicted** the leading claim in the literature, which predicted that only certain "untopicalizable" adverbs would sound off. The takeaway was a methodological one too: corpus data alone (looking at what people *write*) isn't enough to draw conclusions about grammar — you need judgment experiments to test what people actually *accept*.

---

# Paper Abstract

Putting Theories of Adverb Topicalization in Irregular Wh-Questions to the (Acceptability
Judgement) Test
Testing Theories of Adverb Topicalization in Irregular Wh-Questions
Judgment Testing of Adverb Topicalization in Irregular Wh-Questions
Background: In current scholarship, irregular wh-questions (IWQs)—wh-questions in
English such as _what about_ —are used as diagnostic tools in discourse studies, in which IWQs are
followed by a gerund clause or NP (Huddleston et al. 2002). However, recent data extracted from
the British National Corpus (BNC) led Li & Lui (2023) to purport that “topicalizable” adverbs like
domain, locative, and temporal adverbs can follow IWQs, but “untopicalizable” adverbs like
epistemological, predicational, and discourse-oriented adverbs cannot. The current study tests Li
and Lui’s (2023) claim through a more rigorous method of examination—a contextualized
acceptability judgment task (AJT)—to explore the notion that relying on corpora alone to study
syntax is problematic.
Method: Data was collected via a 7-point Likert scale AJT (12 contextualized tokens and 24
contextualized fillers) from 51 native speakers of American English. To examine the interaction
between IWQ and adverb type, the experiment utilized a 2x2 factorial design that crossed two
factors: question type (irregular wh-question vs. wh-question) and adverb type (untopicalizable
adverb vs. topicalizable adverb). Table 1 represents the test design with a sample token set.
_Table 1: 2x2 Factorial Design_
Context: Tim loved to listen to music in the car, but his little brother, a talented piano player,
constantly critiqued his taste and asked annoying questions about aspects of the music. This
morning, Tim’s brother asked:
Question Type
Adverb Type
**Irregular Wh-Question Wh-Question
Untopicalizable** (1) * What about clearly the song’s sound? (2) What is clearly the song’s sound?
**Topicalizable** (3) What about tonally the song’s sound? (4) What is tonally the song’s sound?
Results: We conducted a linear mixed-effects regression analysis on z-transformed ratings with
question and adverb types included as fixed factors and subject and item included as random
factors. Results indicated significant effects of question type (Estimate = 0.267, SE = 0.132, p <
0.05), as wh-questions were rated higher than irregular wh-questions. However, there was no
significant interaction effect of question and adverb types, indicating that participants rated irregular
wh-questions lower, regardless of adverb type. This contradicts the claim made by Li & Lui (2023).
Discussion and Conclusions: A contextualized AJT failed to find the interaction between
question and adverb types purported by Li & Lui’s (2023) probing of the BNC. This underscores the


potential pitfalls of using production data like corpora to study syntax because “absence of evidence
does not mean evidence of absence”; in other words, merely because IWQs with untopicalizable
adverbs are not frequently found in the BNC does not mean that they are necessarily
ungrammatical. Nonetheless, the interaction between question and adverb types warrants further
research: If adverbials can appear as a topic in IWQs as Li & Lui (2023) claim—a novel assertion
that adds to the existing literature on topics and adverbials—then that fact should be corroborated
by judgment data.
**References:**
Huddleston, R., et al. (2002). _The Cambridge grammar of the English language_. Cambridge:
Cambridge University Press.
Li, W., & Liu, J. (2023). About ‘what about’: the semantics and syntax of irregular wh-questions in
English. _Linguistics_ 61(1). 159–195.



# Paper 3-4 page summary (original was 27 pages):

---

**Summary: "Irregular Wh-Questions and Adverbs in English: A Syntactic Examination"**
*Boes, Klein, and Wacker — English 420: Experimental Syntax, UW-Madison (2023)*

---

**Introduction**

English speakers use phrases like *"What about tomorrow?"* or *"What about politically?"* without thinking twice — but these constructions, known as **irregular wh-questions (IWQs)**, have received surprisingly little formal attention in syntactic theory. IWQs, defined as wh-questions formed with *what* or *how* followed by *about* or the conditional *if*, are typically treated in the literature merely as conversational tools — used to issue directives, introduce new topics, or make suggestions (Quirk et al. 1987; Huddleston et al. 2002). Crucially, existing descriptions assume IWQs are followed by a noun phrase, leaving their interaction with adverbials largely unexamined.

That gap began to close with Li and Lui (2023), who mined the British National Corpus (BNC) and found that certain adverbs *can* follow *what about* in IWQs — but not all of them. Specifically, **topicalizable adverbs** like domain (*politically*), locative (*upstairs*), and temporal (*tomorrow*) adverbs appear grammatically after *what about*, while **untopicalizable adverbs** like epistemological (*possibly*), predicational (*honestly*), and discourse-oriented (*similarly*) adverbs do not. Li and Lui concluded that IWQs take a wider range of forms than previously described, but their findings rested entirely on corpus data — meaning they observed what speakers *produce*, not what they *accept* as grammatical.

This study steps in to fill that gap. It is the **first to experimentally test** whether the topicalizability distinction identified by Li and Lui actually holds up under scrutiny — using a **contextualized acceptability judgment task** administered to native English speakers. In doing so, it situates the study of IWQs within the framework of experimental syntax, placing questions of discourse and grammaticality in direct conversation with each other.

---

**Literature Review**

To understand why adverbs behave differently in IWQs, the paper surveys two dominant theories of adverbial distribution in syntax: the **cartographic theory** and the **scopal theory**.

The cartographic theory, associated with Cinque (1999), holds that syntax is the primary driver of where adverbs can appear. Under this view, each adverb is licensed by a dedicated functional head in a rigidly ordered hierarchy dictated by Universal Grammar — meaning every adverb has a fixed syntactic home. The scopal theory, advanced by Ernst (2002), takes the opposite tack: adverbs are governed by semantics, and can appear anywhere in a sentence as long as doing so doesn't produce a semantic violation. Both theories have well-documented shortcomings — the cartographic theory struggles to account for cases where the same position seems to carry different meanings, while the scopal theory incorrectly predicts free ordering in cases where adverb order is actually quite rigid.

More importantly for this study, **both theories fail to account for adverbials in IWQs**. The cartographic theory assumes adverbials occupy either a specifier position within a functional projection or the complement position of a verb phrase — but IWQs contain no explicit verb, and adverbs of different types appear in the same syntactic position within them, violating both assumptions. The scopal theory, meanwhile, predicts that different adverbials occupy different syntactic positions, yet in IWQs all adverbials appear where they are base-generated, and both arguments and adverbials surface in the same position — a fact neither theory can accommodate.

This theoretical vacuum is partly addressed by **Ernst (2020)**, who draws a distinction between *topicalizable* and *untopicalizable* modifiers. Domain adverbs, he argues, can function as topics that represent sets of propositions — much like temporal adverbials — while predicational adverbs cannot be topicalized. His examples illustrate this cleanly: *"Politically, why would this be a problem?"* is natural, while *"Oddly, why would this be a problem?"* is not.

The semantic explanation for this distinction comes from **Jacobs (2001)**, whose theory of topic construction identifies **frame-setting** as the key attribute at play in IWQs. A frame, in Jacobs's terms, is a dimension of a topic that restricts the proposition in the rest of the sentence to a particular domain of reality. Domain, locative, and temporal adverbs all specify such a domain clearly — *politically* restricts a proposition to the political sphere, *tomorrow* to a specific time — making them natural fits as IWQ complements. Epistemological, predicational, and discourse-oriented adverbs, by contrast, fail to specify any such domain, which is why sentences like *"What about possibly?"* strike speakers as odd. The theoretical prediction, then, is clear: topicalizability is at least partly a semantic property, and adverb type should meaningfully impact grammaticality judgments in IWQs. This study sets out to test that prediction empirically for the first time.

---

**Methodology**

The central research question of the study asks how adverb type and question type interact to affect participants' grammaticality ratings in an acceptability judgment task. More specifically, the study narrows its focus to **two untopicalizable adverb types** — epistemological and predicational — and **one topicalizable type** — domain adverbs — tested across both irregular and regular wh-questions. The hypothesis was that IWQs paired with untopicalizable adverbs would receive the lowest ratings, while IWQs with topicalizable adverbs and regular wh-questions with either adverb type would be rated more favorably.

Participants were 20 native English-speaking undergraduates at UW-Madison between the ages of 18 and 22, balanced for gender. Linguistics students were excluded to prevent prior knowledge from influencing responses.

The experiment used a **2x2 factorial design**, crossing two factors — question type (irregular vs. regular wh-question) and adverb type (topicalizable vs. untopicalizable) — each at two levels, yielding four conditions per token set. Twelve token sets were developed (with plans for sixteen total), each consisting of four lexically matched sentences varying only in question type and adverb type. For example, one token set presented the context of a radiologist asking about a patient's arm, and tested sentences like *"What about specifically the position of your arm?"* (IWQ + untopicalizable) against *"What about physically the position of your arm?"* (IWQ + topicalizable), and their regular wh-question counterparts.

To control for bias, 32 contextualized filler items were included — 12 grammatical and 20 ungrammatical — ensuring test items comprised only about 30% of each list. Additional fillers beginning with *what* were added to prevent participants from associating the word with grammaticality. The items were arranged into **four test lists using a Latin Square design** and pseudorandomized to prevent test items from appearing adjacently. Six practice items opened each list to familiarize participants with the rating scale.

Responses were collected via an online survey (Google Forms or Qualtrics) using a **7-point Likert scale**, with 7 indicating fully acceptable and 1 indicating fully unacceptable. To correct for individual differences in scale usage, all ratings were **z-transformed** before analysis.

---

**Results and Discussion**

After removing outliers — participants whose rating standard deviation fell more than one standard deviation outside the group mean — the data were analyzed in R using a linear mixed-effects model. The two crossed fixed factors were question type and topicalizability; random effects included age and gender, neither of which showed significant variance.

Descriptively, the mean ratings across all four conditions were notably close together and low: irregular wh-questions with topicalizable adverbs averaged 2.615, and with untopicalizable adverbs 2.583 — nearly identical. Regular wh-questions fared only slightly better, with topicalizable adverbs averaging 3.066 and untopicalizable 2.971. For reference, grammatical fillers averaged 4.423 and ungrammatical fillers 2.271 — suggesting all four experimental conditions were clustering near the ungrammatical end of the scale, with high standard deviations (all above 20%) indicating considerable variability in responses.

The mixed-effects model found a **significant effect of question type** (estimate = 0.37, p = 0.03): regular wh-questions were rated higher than irregular wh-questions overall. However, the predicted **interaction between question type and adverb type was not significant** — meaning participants rated irregular wh-questions poorly regardless of whether the adverb was topicalizable or not. A comparison of the two untopicalizable adverb types (epistemological: mean 2.703; predicational: mean 2.877) also showed no significant difference between them.

These results offer **partial support** for the original hypothesis. The prediction that untopicalizable adverbs would fare worse than topicalizable ones held for regular wh-questions, but crucially failed for IWQs — where both adverb types were rated equally poorly. This is a meaningful finding: rather than confirming Li and Lui's (2023) claim that only untopicalizable adverbs produce ungrammatical IWQs, the data suggest that **any adverb paired with an irregular wh-question tends to be rated as unacceptable**, topicalizable or not.

The authors are candid about the study's limitations. High p-values throughout and wide standard deviations point to significant variability, likely stemming from small sample size and design factors that could be refined. Still, the directional trends are suggestive, and the authors argue that the interaction between question type and adverb type warrants further investigation with a more powered study. More broadly, the paper makes a methodological point: corpus data, which records only what speakers *produce*, cannot substitute for judgment data that tests what speakers *accept*. The absence of untopicalizable adverbs in IWQs in the BNC does not mean they are ungrammatical — and experimental syntax is the right tool to sort that out.
