Self-nonself discrimination in the immune system
Our immune systems play a vital role in defending us against pathogens. However, in order to do its job, the immune system has to perform the nontrivial task of distinguishing—at a microscopic level—between molecules that belong to its host organism, and foreign (potentially pathogenic) molecules that indicate infection. This is a difficult task because all of biology (as we know it) is made of similar stuff: the amino acids, nucleotides, sugars, etc of which life is composed are shared across the whole tree of life: what lends biology its diversity is how these building blocks are put together. Therefore, if one is looking only at the building blocks, which the immune system is in many cases constrained to do, it can be hard to tell self from nonself, just as telling your home apart from that of your neighbor would be difficult if you only had access to the bricks and mortar that were used to construct them.
The question of the how the immune system (here focusing on the adaptive immune system) distinguishes self from nonself is a classical question in immunology, but still remains incompletely answered. Here we will discuss a quantitative approach: essentially a consistency check showing that empirical numbers from the immune system are consistent with the simplest possible model of interaction between immune receptors and antigens (the stuff which the immune cells evaluate as either self or not). We will use empirical numbers from humans, whose immune systems have been intensely studied. The ideas herein come from a 1993 paper by Percus, Percus and Perelson; after discussing their argument we will then extend their ideas to incorporate a mechanism of central tolerance (the process by which immune systems are trained to avoid attacking self-antigens) that is known to decrease the prevalence of autoimmunity.
The task of distinguishing self from nonself
The recognition of foreign antigens in our bodies is carried out in part by our body’s assemblage of B cells and T cells. These cells possess receptors (B cell receptors and T cell receptors or BCRs and TCRs respectively) on their surface: these receptors query their surroundings, and when they detect foreign antigens (in the case of T cells, presented to them via special molecules on the surface of other cells) they have the capacity to mount an immune response. This mechanism raise two questions: (1) how are receptors on immune cells able to query the identity of the huge diversity of possible antigens found in the body? and (2) How do the receptors distinguish between self and nonself antigens?
The answer to (1) is that our body is home to a diverse library of both BCRs and TCRs. This library is composed of \(\sim 10^7\) clonotypes (unique receptor identities), whose combined binding coverage of the space of antigens presumably allows at least one receptor to bind to at least some part of any possible pathogen. There is no way that all this diversity could be genetically encoded: our genomes are far too short. Instead it is generated through a stochastic process of genetic recombination, in which groups of genes are shuffled around (largely in the fetus and young infants) in order to give rise to a random sample of receptors that is unique to each individual, since it vastly undersamples the possible combinatorial diversity that it could generate.
The answer to (2) is that we are under strong selection pressure to evolve immune systems that do not classify self antigens as foreign: failure to achieve this classification would trigger an autoimmune response. The structure of our immune repertoires is therefore a product both of selection pressure on the genes involved in receptor diversity, and also of dynamical processes that occur during our lifetime, wherein our immune cells undergo training or “education” to avoid recognizing self antigens as foreign. There is also a major role played by regulatory immune cells which tamp down excessive self-reactivity, but which we will not discuss much here. We will first discuss properties of the immune repertoire that might have emerged over generations of evolution, and then look at how these properties would be modified by an active “training” of the immune system that is known to happen during an individual’s lifetime.
Properties of an evolved immune system
Percus, Percus and Perelson were interested in making quantitative the question of self-nonself discrimination. In particular: are empirically observed properties of receptor-antigen binding consistent with the simplest nontrivial model in which every receptor binds every antigen (both self and nonself) independently? In such a model there is a microscopic probability, \(p\), of each receptor binding to each antigen, that maximizes the macroscopic probability that the immune repertoire will successfully distinguish self and nonself. This optimal \(p\) sets a scale for a typical antigen-receptor affinity towards which evolution might have driven our immune systems.
We consider a model of the immune system whose job it is to distinguish between \(N\) self antigens and \(N'\) foreign antigens with a repertoire of \(n\) receptors. These could be either T cell or B cell receptors; for now we remain intentionally vague. Successful discrimination means that at least one receptor in the repertoire binds to each foreign antigen (and can then trigger an immune response), and that none of the receptors bind to any of the self antigens. Let us define \(q\equiv 1-p\) as the probability of a random receptor not binding to a random antigen: then the probability that a single foreign antigen is bound by at least one of the receptors in a repertoire of size \(n\) is \(1-q^n\), and the probability that a given self antigen is not bound by any of the receptors in the repertoire is \(q^n\). Therefore, under the assumption that each antigen is independent (surely an inaccurate assumption, but perhaps a reasonable starting point), the probability that our randomly assembled repertoire successfully discriminates between self and nonself is
\[P_{\it dis} = (1-q^n)^{N'} q^{nN}.\]The argument made by Percus, Percus and Perelson is that evolution should tend to drive \(q\) to the value that maximizes \(P_{\it dis}\). One can see that for \(q\) close to \(1\) (i.e. very low probability of receptor-antigen binding) the probability of having a functional repertoire is very small: this is because it becomes unlikely to have every foreign antigen covered by at least one receptor. At the same time as \(q\) get close to \(0\), the probability of having a functional repertoire becomes small because it is unlikely to have the whole repertoire avoid binding to each foreign antigen. Taking the derivative of \(P_{\it dis}\) with respect to \(q\), setting to \(0\), and solving for \(q\), gives
\[q^* = \left(1 + \frac{N'}{N}\right)^{-1/n} \approx 1 - \frac{1}{n} \log\left(1+\frac{N'}{N}\right).\]Recalling that \(q\) is the probability of a random receptor not binding to an antigen, we can see that the optimal binding probability is
\[p^* \approx \frac{1}{n} \log\left(1+\frac{N'}{N}\right).\]Empirical numbers
We would like to check whether these numbers are consistent with empirical data. What are reasonable estimates of the numbers \(n\), \(N\), and \(N'\)? As stated before, the number of unique receptors (or clonotypes) in an individual’s repertoire is roughly \(n\sim 10^7\). The number of self antigens in an individual can be estimated from the number of proteins in the human genome, which is \(N\sim 10^{4.5}\). The number of foreign antigens is difficult to estimate, but one way to do so involves considering the total possible diversity of the immune repertoire, and reasoning that each possible receptor that could be produced should bind to at least some epitope: therefore the number of foreign epitopes is similar to the amount of possible receptor diversity (much larger than the repertoire diversity in a single person). An estimate of the total amount of receptor diversity that can be generated through VDJ recombination gives \(N' \sim 10^{16}\).
Plugging in these three numbers gives \(p^* \approx 10^{-6}\). This estimate is within an order of magnitude of experimental results in which the number of peptides (for our purposes these are the relevant antigens) that binds to a single TCR was measured Wooldridge et al 2012. The fraction of peptides that binds to a single TCR (averaged over TCRs) is equal to the fraction of TCRs that binds to a single antigen (averaged over antigens), and is therefore a measurement of \(p^*\).
Note that, even with \(q=q^*\), i.e. for the optimal microscopic binding probability, there is a tiny probability of the immune system actually correctly distinguishing self from nonself by pure chance (given reasonable numbers for the human immune system). However, we can ascribe some meaning to our result for \(q^*\) as follows: say that we are told that a particular biased coin, when flipped 100 times, yielded 30 heads and 70 tails. Our best guess for the heads-probability of the coin is 30%. Despite the fact that the probability of getting exactly 30 heads with a heads-probability of 30% is very small, we should expect that the heads-probability of the coin is in the ballpark of 30%. In the same way, even if the probability of a successfully discriminatory repertoire is very small under the model of independent receptor-antigen binding for every pair, we aim to find the \(p\) (and therefore the \(q\)) that maximizes the probability of having such a repertoire. Indeed the real immune repertoire has neither independent antigen-receptor binding probabilities nor perfectly differentiates self from nonself. The former means that a more realistic model would have some structure in these binding interactions, and the latter allows us some leeway in specifying that the immune repertoire distinguish self from nonself—the decision boundary it computes can be fuzzy.
Epitope size
Percus, Percus and Perelson proceed to translate the binding probability between receptor and antigen into a characteristic size of the portion that a receptor recognizes. These portions of antigens to which receptors bind are called epitopes: for B cells they tend to be pieces of e.g. a protein whose physical shape is recognized by the BCR expressed on the B cell surface, while for T cells, the relevant epitopes are MHC-peptide complexes, where the peptide is a fragment of a protein either expressed in the human proteome or found in the body due to infection or foreign invasion. From the optimal binding probability one can argue—under the model of binding being determined by a sufficient number of matches between epitope and receptor—for an epitope size that is consistent with this binding probability, where the epitope size is essentially the number of contiguous matches that ones needs to require between the antigen and receptor such that binding to the antigen happens with the optimal probability \(p^*\).
Thymic selection
What happens if we try to extend the argument above to an immune system that is not entirely randomly assembled, but which undergoes some form of training, known as central tolerance? We will try to model the process of thymic selection, wherein T cells are presented with some subset of human antigens, and are removed from the immune repertoire if their reactivity to a self antigen that they encounter is too high. Our model will be of a randomly assembled immune repertoire (as before) for which each receptor is presented with \(s\) self antigens, randomly chosen from the set of self antigens. If the receptor binds to any of these self antigens, it will be eliminated from the repertoire. B cells undergo a more lenient process of central tolerance in which their BCRs can be edited if they bind too strongly to self antigens, but below we will keep the TCR central tolerance mechanism as our motivating example.
What, then, is the probability of our immune repertoire successfully discriminating self from nonself? Note that \(n\) is the size of our repertoire before thymic selection. After thymic selection, only the receptors that did not bind to any of the presented self antigens (where we assume each receptor is presented with a random subset of \(s\) self antigens) will remain. Therefore the post-selection size of the immune repoertoire is \(q^sn\). However, thymic selection not only reduces the diversity of the repertoire: it also decreases the number of self antigens that the repertoire needs to avoid: in particular each receptor is now guaranteed to not bind to a random subset of \(s\) self antigens. Therefore the probability that a random post-selection receptor does not bind to any self antigens is \(q^{N-s}\). As a result we can write the following expression for the probability of successful distinguishing between self and nonself:
\[P_{\it dis} = (1-q^{q^sn})^{N'} q^{(N-s)q^s n}.\]Now we will optimize over both \(s\) and \(q\): taking both these derivatives of \(P_{\it dis}\) and setting them to zero gives two equations:
\[s = -\frac{1}{\log q},\quad \frac{N}{N'} = \frac{q^{q^s n}}{1-q^{q^s n}}\]which have solution
\[q = \left(1 + \frac{N'}{N}\right)^{-e/n},\quad s = \frac{n}{e} \log^{-1}\left(1+\frac{N'}{N}\right),\]where we used the fact that \(q^s = 1/e\), which indicates that the process of “optimal” thymic selection reduces repertoire size by the constant factor \(1/e\), irrespective of repertoire size. Note that \(s\) cannot be larger than \(N\). Following the approximation of Percus, Percus and Perelson we see that
\[p^* \approx \frac{e}{n} \log\left( 1 + \frac{N'}{N}\right),\quad s^* \approx \frac{1}{p^* }.\]We therefore see that, compared to the case without thymic selection, the optimal binding probability increases by a constant factor (irrespective of library size) and the optimal number of self antigens to train on is the number such that there is one binding event expected between the TCR under consideration and the antigens it is presented with. Because \(p^*\) depends weakly on \(N\) and \(N'\), so does \(s^*\): therefore for a smaller number of self-antigens, one should ideally show a larger fraction of them to each T cell during thymic education.
In reality, the number of self-peptides shown to a TCR during central tolerance is thought to be one or two orders of magnitude less than \(1/p^*\) as quoted above. In fact, the immune system makes use of other mechanisms of tolerance which make central tolerance only one piece in the puzzle, as we discuss below.
Other mechanisms of self tolerance
So far we se have seen how thymic selection, even though it gets rid of some T cell diversity that might be useful for binding to foreign antigens, is helpful because it builds the constraint of not binding to self antigens into the post-selection repertoire. However there are other mechanisms (both conjectured and known) that are crucial for making the immune system tolerate self antigens. One such mechanisms is quorum sensing: the idea that by making the decision to mount an immune response a collective one between multiple T cells, one can moderate the effect of self-reactive TCRs. Another mechanism is the presence of regulatory T cells which prevent immune response against self antigens that they bind strongly to.
In general, it remains a question what kind of structure in receptor-antigen binding space is conducive to an immune system that can effectively distinguish between self and nonself.