14 Where next
Every model in this book can be fitted with a package today. This chapter lists what those models leave out and points to recent work on each gap, in typology and in the statistics it borrows from.
This chapter has no code. Each section names a limit of the models in the earlier chapters and then lists recent papers that work on it. The entries are short. They say what a paper is about and which chapter it speaks to, and they leave the findings to the papers.
14.1 The assumptions this book made
Five choices were made early and kept throughout.
- One tree, taken as known. Chapter 2 built it from a classification and compared it with one dated tree. No model in the book carried the uncertainty of the tree into the slope.
- One point per language. Chapter 5 showed how far a point can sit from the area where a language is spoken, and how much a location has moved between the traditional and the contemporary record.
- Contact as a smooth surface or a fixed graph. The Gaussian process of Chapter 6 and the neighbour graph of Chapter 7 both say that nearby languages are alike. Neither says who borrowed from whom, when, or across which barrier.
- One slope for the whole world. Elevation had the same effect in the Andes and in the Caucasus.
- One outcome at a time. Chapter 11 changed the family, and each model still had a single feature on its left-hand side.
The sections below come back to each of these, after a look at how the field is using the models that already exist.
14.2 Typology with both controls
The models of Chapters 9 and 10 are now common in quantitative typology, and part of the discussion has moved to how results hold up under them.
- Replication. Becker and Guzmán Naranjo (2025) is a target article on replication and methodological robustness in quantitative typology. The same issue of Linguistic Typology carries commentaries, among them Mauri and Sansò (2025) on replicability “all the way up” and Coupé (2025) on what counts as advanced statistical modelling.
- Results that change under the controls. Blum (2025) reports that most of the over-representation of phonological features in basic vocabulary disappears once spatial and phylogenetic dependence are controlled for. This is the pattern of Chapter 1 on a different question.
- Spatial effects as the object of study. Hartmann and Nichols (2025) model geospatial effects on phonological complexity across the world’s languages, and Di Garbo and Kapellis (2025) survey contact effects in nominal number systems worldwide. In both, geography or contact is the object of study.
- Environment and language. The running example of this book is one of many claims that link a linguistic feature to the physical environment. Robbeets and Heggarty (2026) set out first principles for work on language, ecology and climate change.
- The same problem elsewhere. Wyner and Brill (2026) correct for spatial and temporal dependence in estimates of the economic effect of extreme heat. The data differ and the problem is the one in Chapter 1.
14.3 Sampling
Chapter 12 treated a balanced sample as a statistical method and Chapter 13 compared it with the models. Work on sampling continues alongside the modelling.
- Miestamo (2025) proposes phenomenon-based sampling. Compare it with the variety and probability samples of Chapter 12.
- Villa (2025) is a Python library that draws typological samples with the diversity value metric. A sample drawn by code can be redrawn many times, which is what Chapter 12 asks for.
14.4 Better trees, and more than one
The covariance matrix A of Chapter 3 is only as good as the tree behind it. Chapter 2 listed the sources. The papers below are about how those trees are inferred and how sure anyone can be of them.
- The data behind lexical trees. Häuser and Stamatakis (2025) describe a cognate data bottleneck in language phylogenetics, and Häuser et al. (2026) explore the limits of inference from cognate data. Snee et al. (2026) trace variation between language phylogenies to variation in how the concepts of a word list were translated.
- New family trees. Koile et al. (2025) reconstruct the Semitic tree with sampled ancestors, which lets attested ancient languages sit on the branches as ancestors of later ones. Jacques et al. (2025) use Bayesian phylogenies of Kra-Dai to study the history of loom technology together with the languages.
- Other sources of distance. Mavridis et al. (2026) build phonological distances for typology and apply them to the origin of Indo-European.
- Software. Baele et al. (2025) present BEAST X, the current release of the software behind most dated language trees. Swanepoel et al. (2026) describe TreeFlow, which brings automatic differentiation to phylogenetic models.
- Uncertainty about the tree. Klawitter and Drummond (2026) define credible sets for tree topologies. Piccoli et al. (2025) ask whether comparative models should be weighted when they account for phylogenetic uncertainty.
That last question has a practical version in this book’s terms. A posterior sample of trees gives a sample of matrices A, one per tree. brms::brm_multiple() takes a list of data sets and a matching list of data2 objects. Repeat the data once per tree and supply one matrix per tree, and it fits the model once per matrix and pools the draws. The slope then carries the uncertainty of the tree. Most Phlorest datasets include a posterior sample of trees beside the summary tree used in Chapter 2, so this needs no new method, only time.
14.5 Contact as a process
A Gaussian process says that similarity decays with distance. It has no borrowing events, no direction and no dates. Several lines of work model contact itself.
- Contact inside the history of a family. Santos et al. (2026) model the prehistory of Bantu with coalescent theory and treat contact as a normal part of that history.
- Evidence from outside linguistics. Graff et al. (2025) use patterns of genetic admixture to identify contact and compare rates of borrowing across contact scenarios. Fehn et al. (2025) trace contact and migration in southern Africa before the Bantu expansion through lexical borrowing.
- Trees or networks. Baraghith (2026) evaluates what network models explain that tree models do not in research on cultural evolution, and Duda (2026) reviews tree thinking across linguistics, archaeology and anthropology.
- Diffusion over a map in time. Wichmann (2025) covers linguistic phylogeography, the inference of where languages were spoken in the past. Burridge and Vaux (2026) infer the dynamics of language change from maps that change over time, and Takahashi et al. (n.d.) combine cellular automata with Bayesian MCMC to recover the history of a diffusion process.
- Areas as discrete units. sBayes (Ranacher et al. 2021) finds contact areas as sets of languages that share features beyond what inheritance and universal preference predict. It answers a question the surface of Chapter 6 cannot ask, namely which languages form an area.
14.6 Phylogenetic regression
The phylogenetic term in this book is a random intercept with Brownian covariance. The comparative-methods literature has kept moving past that.
- Guirguis et al. (2026) argue that a bias they call Occam’s bias undermines inferences from phylogenetic linear models. Read it beside the dissenting estimate in Chapter 4.
- Lau et al. (2026) give an efficient Bayesian method for the joint evolution of a continuous and a discrete trait under a state-dependent Ornstein-Uhlenbeck model. This is a process model of two features evolving together, which a regression with a phylogenetic intercept only approximates.
- Davison et al. (2026) build scalable phylogenetic models that capture trait variation within species. The linguistic counterpart is variation within a language: dialects, or several descriptions of one language, each with its own value.
- Mizuno et al. (2026) give a framework and practical guidance for meta-analysis with both phylogenetic and spatial terms, the same pair of terms as in Chapter 9.
14.7 Spatial models at scale
Five hundred languages are a small spatial dataset. A study of every language in Glottolog, or of dialect data with thousands of sites, runs into the cost of the Gaussian process. Chapter 6 used one approximation and Chapter 8 another. There are more.
- Reviews and benchmarks. Fuentes and Patterson (2026) review fifty years of spatial statistics. Tedesco et al. (2025) benchmark software for spatio-temporal models and include a guide to making R faster. Irawan (2026) benchmarks the INLA-SPDE approach for point processes on parameter recovery and spatial confounding, the two questions of Chapter 13.
- Faster approximations. Song and Datta (2026) give a fast variational method for large spatial data. Fu et al. (2026) extend nearest-neighbour processes beyond the Gaussian case. Gaedke-Merzhäuser et al. (2025) accelerate spatio-temporal models with several Gaussian processes on high-performance hardware. Aiello and Banerjee (2026) use amortized inference, in which a model is trained once and then reused, for disease mapping and boundary detection on spatial graphs. Boundary detection is the statistical name for finding an isogloss.
- INLA. Dutta et al. (2026) work on skewness in the INLA approach with variational Bayes. Chapter 8 found a binary model where INLA’s default mode and its classic mode disagreed. Posteriors for binary data are often skewed, so read this before trusting either mode on such data.
- Packages. Finley (2026) is an R package for Bayesian spatial and space-time linear mixed models.
14.8 What the spatial term does to the slope
Chapter 9 warned that a spatial term can move a slope when the predictor is itself smooth in space, and elevation is. This is spatial confounding, and the literature on it is not settled.
- Lin and Warren (2026) ask when spatial random effects are needed at all in Bayesian regression for multilevel areal data.
- Akbari et al. (n.d.) review spatial causality, and He et al. (2026) estimate treatment effects that vary over space with a causal forest on a graph.
- Yacine and Thomas (2026) fit spatially varying coefficients with a Vecchia approximation. A spatially varying coefficient is a slope that changes over the map. For this book’s example it would let elevation matter in one highland region and not in another.
brms can already fit one. The term gp(x, y, by = elev_km) multiplies a spatial surface by elevation, which gives a slope that varies over the map. It needs more data than 50 languages with ejectives can supply, which is why it does not appear in the earlier chapters.
14.9 Who is a neighbour
Chapter 7 joined each language to its five nearest neighbours and checked the result against ten. The graph was fixed before the model saw the data. A group of recent papers estimates it.
- Mendez (n.d.) and Krisztin and Piribauer (2026) estimate the spatial weight matrix inside a Bayesian model. The second is an R package.
- Arumningtyas et al. (2026) compare spatial matrices for the BYM model, the model of Chapter 7.
- Schenk and Afkham (2026) infer the weights of a graph from dynamics inspired by the SPDE approach of Chapter 8.
For languages this is the natural next step. A neighbour graph built from distance treats a mountain range and a trade route alike. A graph with estimated weights could find that two languages on opposite sides of a strait behave as neighbours and two on opposite sides of a ridge do not.
14.10 Languages as areas
Chapter 5 drew speaker areas as polygons and then went back to points, because every model in the book wants one location per language. Two recent papers work with the areas themselves.
-
Cunha Godoy et al. (2026) define a Gaussian process over areal units with the Hausdorff distance, a distance between two shapes, in place of the distance between two points.
sf::st_distance()computes it withwhich = "Hausdorff", so the polygons of Chapter 5 can supply the distances. - Villejo et al. (2026) model outcomes that were recorded for areas as coming from a surface that is continuous in space.
The maps have a history of their own. Jagessar (2025) examines Grierson’s incomplete linguistic map of India and the limits of Indo-European cartography. Read it with the last section of Chapter 5.
14.11 More kinds of outcome
Chapter 11 covered a measurement, a count and ordered classes. The spatial literature has models for cases it left out.
- Counts. Nadifar et al. (2026) model dispersed counts with INLA. Majumder et al. (2026) and Wu et al. (2026) treat zero-inflated negative binomial counts in space and time.
- Discrete data in general. Carter (2024) is a thesis on Bayesian spatial models for discrete data.
- Coding error. Ma et al. (2026) analyse binary spatial data that are sometimes misclassified. Typological codings are too. A grammar that does not mention a feature is often coded as lacking it, and a model that allows for misclassification is the principled way to say so.
- Many features at once. Mukherjee et al. (2026) give scalable inference for high-dimensional spatial data of mixed types. A typological database is this kind of data: hundreds of features, some binary, some ordered, some counts, on one set of locations.
14.12 Predicting what is missing
The dependence terms in this book were used to protect a slope. They are also prediction tools, since a model that knows a language’s relatives and neighbours can guess its missing values. Work on filling typological databases is growing, and much of it comes from natural language processing. Ioannidou (2025) predicts Grambank features, Hus and Anastasopoulos (2026) complete typological databases with retrieval-augmented generation, and Wang et al. (2026) use in-context learning with large language models. A phylogenetic and spatial model is the baseline such methods should be compared with, because it uses only the tree and the map.