# Dirichlet Process. Yee Whye Teh, University College London

Save this PDF as:

Size: px
Start display at page:

## Transcription

1 Dirichlet Process Yee Whye Teh, University College London Related keywords: Bayesian nonparametrics, stochastic processes, clustering, infinite mixture model, Blackwell-MacQueen urn scheme, Chinese restaurant process. Definition The Dirichlet process is a stochastic proces used in Bayesian nonparametric models of data, particularly in Dirichlet process mixture models (also known as infinite mixture models). It is a distribution over distributions, i.e. each draw from a Dirichlet process is itself a distribution. It is called a Dirichlet process because it has Dirichlet distributed finite dimensional marginal distributions, just as the Gaussian process, another popular stochastic process used for Bayesian nonparametric regression, has Gaussian distributed finite dimensional marginal distributions. Distributions drawn from a Dirichlet process are discrete, but cannot be described using a finite number of parameters, thus the classification as a nonparametric model. Motivation and Background Probabilistic models are used throughout machine learning to model distributions over observed data. Traditional parametric models using a fixed and finite number of parameters can suffer from over- or under-fitting of data when there is a misfit between the complexity of the model (often expressed in terms of the number of parameters) and the amount of data available. As a result, model selection, or the choice of a model with the right complexity, is often an important issue in parametric modeling. Unfortunately, model selection is an operation that is fraught with difficulties, whether we use cross validation or marginal probabilities as the basis for selection. The Bayesian nonparametric approach is an alternative to parametric modeling and selection. By using a model with an unbounded complexity, underfitting is mitigated, while the Bayesian approach of computing or approximating the full posterior over parameters mitigates overfitting. A general overview of Bayesian nonparametric modeling can be found under its entry in the encyclopedia [?]. Nonparametric models are also motivated philosophically by Bayesian modeling. Typically we assume that we have an underlying and unknown distribu- 1

2 tion which we wish to infer given some observed data. Say we observe x 1,..., x n, with x i F independent and identical draws from the unknown distribution F. A Bayesian would approach this problem by placing a prior over F then computing the posterior over F given data. Traditionally, this prior over distributions is given by a parametric family. But constraining distributions to lie within parametric families limits the scope and type of inferences that can be made. The nonparametric approach instead uses a prior over distributions with wide support, typically the support being the space of all distributions. Given such a large space over which we make our inferences, it is important that posterior computations are tractable. The Dirichlet process is currently one of the most popular Bayesian nonparametric models. It was first formalized in [1] 1 for general Bayesian statistical modeling, as a prior over distributions with wide support yet tractable posteriors. Unfortunately the Dirichlet process is limited by the fact that draws from it are discrete distributions, and generalizations to more general priors did not have tractable posterior inference until the development of MCMC techniques [3, 4]. Since then there has been significant developments in terms of inference algorithms, extensions, theory and applications. In the machine learning community work on Dirichlet processes date back to [5, 6]. Theory The Dirichlet process (DP) is a stochastic process whose sample paths are probability measures with probability one. Stochastic processes are distributions over function spaces, with sample paths being random functions drawn from the distribution. In the case of the DP, it is a distribution over probability measures, which are functions with certain special properties which allow them to be interpreted as distributions over some probability space Θ. Thus draws from a DP can be interpreted as random distributions. For a distribution over probability measures to be a DP, its marginal distributions have to take on a specific form which we shall give below. We assume that the user is familiar with a modicum of measure theory and Dirichlet distributions. Before we proceed to the formal definition, we will first give an intuitive explanation of the DP as an infinite dimensional generalization of Dirichlet distributions. Consider a Bayesian mixture model consisting of K components: π α Dir( α K,..., α K ) θ k H H z i π Mult(π) x i z i, {θ k} F (θ z i ) (1) where π is the mixing proportion, α is the pseudocount hyperparameter of the Dirichlet prior, H is the prior distribution over component parameters θ k, and F (θ) is the component distribution parametrized by θ. It can be shown that for large K, because of the particular way we parametrized the Dirichlet prior over π, the number of components typically used to model n data items 1 Note however that related models in population genetics date back to [2]. 2

3 becomes independent of K and is approximately O(α log n). This implies that the mixture model stays well-defined as K, leading to what is known as an infinite mixture model [5, 6]. This model was first proposed as a way to sidestep the difficult problem of determining the number of components in a mixture, and as a nonparametric alternative to finite mixtures whose size can grow naturally with the number of data items. The more modern definition of this model uses a DP and with the resulting model called a DP mixture model. The DP itself appears as the K limit of the random discrete probability measure K k=1 π kδ θ k, where δ θ is a point mass centred at θ. We will return to the DP mixture towards the end of this entry. Dirichlet Process For a random distribution G to be distributed according to a DP, its marginal distributions have to be Dirichlet distributed [1]. Specifically, let H be a distribution over Θ and α be a positive real number. Then for any finite measurable partition A 1,..., A r of Θ the vector (G(A 1 ),..., G(A r )) is random since G is random. We say G is Dirichlet process distributed with base distribution H and concentration parameter α, written G DP(α, H), if (G(A 1 ),..., G(A r )) Dir(αH(A 1 ),..., αh(a r )) (2) for every finite measurable partition A 1,..., A r of Θ. The parameters H and α play intuitive roles in the definition of the DP. The base distribution is basically the mean of the DP: for any measurable set A Θ, we have E[G(A)] = H(A). On the other hand, the concentration parameter can be understood as an inverse variance: V [G(A)] = H(A)(1 H(A))/(α + 1). The larger α is, the smaller the variance, and the DP will concentrate more of its mass around the mean. The concentration parameter is also called the strength parameter, referring to the strength of the prior when using the DP as a nonparametric prior over distributions in a Bayesian nonparametric model, and the mass parameter, as this prior strength can be measured in units of sample size (or mass) of observations. Also, notice that α and H only appear as their product in the definition (2) of the DP. Some authors thus treat H = αh, as the single (positive measure) parameter of the DP, writing DP( H) instead of DP(α, H). This parametrization can be notationally convenient, but loses the distinct roles α and H play in describing the DP. Since α describes the concentration of mass around the mean of the DP, as α, we will have G(A) H(A) for any measurable A, that is G H weakly or pointwise. However this not equivalent to saying that G H. As we shall see later, draws from a DP will be discrete distributions with probability one, even if H is smooth. Thus G and H need not even be absolutely continuous with respect to each other. This has not stopped some authors from using the DP as a nonparametric relaxation of a parametric model given by H. However, if smoothness is a concern, it is possible to extend the DP by convolving G with kernels so that the resulting random distribution has a density. 3

4 A related issue to the above is the coverage of the DP within the class of all distributions over Θ. We already noted that samples from the DP are discrete, thus the set of distributions with positive probability under the DP is small. However it turns out that this set is also large in a different sense: if the topological support of H (the smallest closed set S in Θ with H(S) = 1) is all of Θ, then any distribution over Θ can be approximated arbitrarily accurately in the weak or pointwise sense by a sequence of draws from DP(α, H). This property has consequence in the consistency of DPs discussed later. For all but the simplest probability spaces, the number of measurable partitions in the definition (2) of the DP can be uncountably large. The natural question to ask here is whether objects satisfying such a large number of conditions as (2) can exist. There are a number of approaches to establish existence. [1] noted that the conditions (2) are consistent with each other, and made use of Kolmogorov s consistency theorem to show that a distribution over functions from the measurable subsets of Θ to [0, 1] exists satisfying (2) for all finite measurable partitions of Θ. However it turns out that this construction does not necessarily guarantee a distribution over probability measures. [1] also provided a construction of the DP by normalizing a gamma process. In a later section we will see that the predictive distributions of the DP are related to the Blackwell-MacQueen urn scheme. [7] made use of this, along with de Finetti s theorem on exchangeable sequences, to prove existence of the DP. All the above methods made use of powerful and general mathematical machinery to establish existence, and often require regularity assumptions on H and Θ to apply these machinery. In a later section, we describe a stick-breaking construction of the DP due to [8], which is a direct and elegant construction of the DP which need not impose such regularity assumptions. Posterior Distribution Let G DP(α, H). Since G is a (random) distribution, we can in turn draw samples from G itself. Let θ 1,..., θ n be a sequence of independent draws from G. Note that the θ i s take values in Θ since G is a distribution over Θ. We are interested in the posterior distribution of G given observed values of θ 1,..., θ n. Let A 1,..., A r be a finite measurable partition of Θ, and let n k = #{i : θ i A k } be the number of observed values in A k. By (2) and the conjugacy between the Dirichlet and the multinomial distributions, we have: (G(A 1 ),..., G(A r )) θ 1,..., θ n Dir(αH(A 1 ) + n 1,..., αh(a r ) + n r ) (3) Since the above is true for all finite measurable partitions, the posterior distribution over G must be a DP as well. A little algebra shows that the posterior DP has updated concentration parameter α + n and base distribution αh+ n i=1 δ θ i α+n, where δ i is a point mass located at θ i and n k = n i=1 δ i(a k ). In other words, the DP provides a conjugate family of priors over distributions that is closed under posterior updates given observations. Rewriting the posterior DP, we 4

5 have: ( G θ 1,..., θ n DP α + n, α α+n H + n α+n n ) i=1 δ θ i n (4) Notice that the posterior base distribution is a weighted average between the n i=1 prior base distribution H and the empirical distribution δ θ i n. The weight associated with the prior base distribution is proportional to α, while the empirical distribution has weight proportional to the number of observations n. Thus we can interpret α as the strength or mass associated with the prior. In the next section we will see that the posterior base distribution is also the predictive distribution of θ n+1 given θ 1,..., θ n. Taking α 0, the prior becomes non-informative in the sense that the predictive distribution is just given by the empirical distribution. On the other hand, as the amount of observations grows large, n α, the posterior is simply dominated by the empirical distribution which is in turn a close approximation of the true underlying distribution. This gives a consistency property of the DP: the posterior DP approaches the true underlying distribution. Predictive Distribution and the Blackwell-MacQueen Urn Scheme Consider again drawing G DP(α, H), and drawing an i.i.d. sequence θ 1, θ 2,... G. Consider the predictive distribution for θ n+1, conditioned on θ 1,..., θ n and with G marginalized out. Since θ n+1 G, θ 1,..., θ n G, for a measurable A Θ, we have P (θ n+1 A θ 1,..., θ n ) = E[G(A) θ 1,..., θ n ] ( ) = 1 n αh(a) + δ θi (A) α + n i=1 (5) where the last step follows from the posterior base distribution of G given the first n observations. Thus with G marginalized out: ( ) θ n+1 θ 1,..., θ n 1 n αh + δ θi (6) α + n Therefore the posterior base distribution given θ 1,..., θ n is also the predictive distribution of θ n+1. The sequence of predictive distributions (6) for θ 1, θ 2,... is called the Blackwell- MacQueen urn scheme [7]. The name stems from a metaphor useful in interpreting (6). Specifically, each value in Θ is a unique color, and draws θ G are balls with the drawn value being the color of the ball. In addition we have an urn containing previously seen balls. In the beginning there are no balls in the urn, and we pick a color drawn from H, i.e. draw θ 1 H, paint a ball with that color, and drop it into the urn. In subsequent steps, say the n + 1st, we α will either, with probability α+n, pick a new color (draw θ n+1 H), paint a n ball with that color and drop the ball into the urn, or, with probability 5 i=1 α+n,

6 reach into the urn to pick a random ball out (draw θ n+1 from the empirical distribution), paint a new ball with the same color and drop both balls back into the urn. The Blackwell-MacQueen urn scheme has been used to show the existence of the DP [7]. Starting from (6), which are perfectly well-defined conditional distributions regardless of the question of the existence of DPs, we can construct a distribution over sequences θ 1, θ 2,... by iteratively drawing each θ i given θ 1,..., θ i 1. For n 1 let P (θ 1,..., θ n ) = n P (θ i θ 1,..., θ i 1 ) (7) i=1 be the joint distribution over the first n observations, where the conditional distributions are given by (6). It is straightforward to verify that this random sequence is infinitely exchangeable. That is, for every n, the probability of generating θ 1,..., θ n using (6), in that order, is equal to the probability of drawing them in any alternative order. More precisely, given any permutation σ on 1,..., n, we have P (θ 1,..., θ n ) = P (θ σ(1),..., θ σ(n) ) (8) Now de Finetti s theorem states that for any infinitely exchangeable sequence θ 1, θ 2,... there is a random distribution G such that the sequence is composed of i.i.d. draws from it: n P (θ 1,..., θ n ) = G(θ i ) dp (G) (9) In our setting, the prior over the random distribution P (G) is precisely the Dirichlet process DP(α, H), thus establishing existence. A salient property of the predictive distribution (6) is that it has point masses located at the previous draws θ 1,..., θ n. A first observation is that with positive probability draws from G will take on the same value, regardless of smoothness of H. This implies that the distribution G itself has point masses. A further observation is that for a long enough sequence of draws from G, the value of any draw will be repeated by another draw, implying that G is composed only of a weighted sum of point masses, i.e. it is a discrete distribution. We will see two sections below that this is indeed the case, and give a simple construction for G called the stick-breaking construction. Before that, we shall investigate the clustering property of the DP. Clustering, Partitions and the Chinese Restaurant Process In addition to the discreteness property of draws from a DP, (6) also implies a clustering property. The discreteness and clustering properties of the DP play crucial roles in the use of DPs for clustering via DP mixture models, described in the application section. For now we assume that H is smooth, so that all i=1 6

7 repeated values are due to the discreteness property of the DP and not due to H itself 2. Since the values of draws are repeated, let θ1,..., θm be the unique values among θ 1,..., θ n, and n k be the number of repeats of θk. The predictive distribution can be equivalently written as: ( ) θ n+1 θ 1,..., θ n 1 m αh + n k δ θ α + n k (10) Notice that value θk will be repeated by θ n+1 with probability proportional to n k, the number of times it has already been observed. The larger n k is, the higher the probability that it will grow. This is a rich-gets-richer phenomenon, where large clusters (a set of θ i s with identical values θk being considered a cluster) grow larger faster. We can delve further into the clustering property of the DP by looking at partitions induced by the clustering. The unique values of θ 1,..., θ n induce a partitioning of the set [n] = {1,..., n} into clusters such that within each cluster, say cluster k, the θ i s take on the same value θk. Given that θ 1,..., θ n are random, this induces a random partition of [n]. This random partition in fact encapsulates all the properties of the DP, and is a very well studied mathematical object in its own right, predating even the DP itself [2, 9, 10]. To see how it encapsulates the DP, we simply invert the generative process. Starting from the distribution over random partitions, we can reconstruct the joint distribution (7) over θ 1,..., θ n, by first drawing a random partition on [n], then for each cluster k in the partition draw a θk H, and finally assign θ i = θk for each i in cluster k. From the joint distribution (7) we can obtain the DP by appealing to de Finetti s theorem. The distribution over partitions is called the Chinese restaurant process (CRP) due to a different metaphor 3. In this metaphor we have a Chinese restaurant with an infinite number of tables, each of which can seat an infinite number of customers. The first customer enters the restaurant and sits at the first table. The second customer enters and decides either to sit with the first customer, or by herself at a new table. In general, the n + 1st customer either joins an already occupied table k with probability proportional to the number n k of customers already sitting there, or sits at a new table with probability proportional to α. Identifying customers with integers 1, 2,... and tables as clusters, after n customers have sat down the tables define a partition of [n] with the distribution over partitions being the same as the one above. The fact that most Chinese restaurants have round tables is an important aspect of the CRP. This is because it does not just define a distribution over partitions of [n], it also defines a distribution over permutations of [n], with each table corresponding to a cycle of the permutation. We do not need to explore this aspect further and refer the interested reader to [9, 10]. This distribution over partitions first appeared in population genetics, where it was found to be a robust distribution over alleles (clusters) among gametes 2 Similar conclusions can be drawn when H has atoms, there is just more bookkeeping. 3 The name was coined by Lester Dubins and Jim Pitman in the early 1980 s [9]. k=1 7

8 (observations) under simplifying assumptions on the population, and is known under the name of Ewens sampling formula [2]. Before moving on we shall consider just one illuminating aspect, specifically the distribution of the number of clusters among n observations. Notice that for i 1, the observation θ i α takes on a new value (thus incrementing m by one) with probability α+i 1 independently of the number of clusters among previous θ s. Thus the number of cluster m has mean and variance: n α E[m n] = = α(ψ(α + n) ψ(α)) α + i 1 i=1 ( α log 1 + n ) for N, α 0, (11) α V [m n] = α(ψ(α + n) ψ(α)) + α 2 (ψ (α + n) ψ (α)) ( α log 1 + n ) for n > α 0, (12) α where ψ( ) is the digamma function. Note that the number of clusters grows only logarithmically in the number of observations. This slow growth of the number of clusters makes sense because of the rich-gets-richer phenomenon: we expect there to be large clusters thus the number of clusters m has to be smaller than the number of observations n. Notice that α controls the number of clusters in a direct manner, with larger α implying a larger number of clusters a priori. This intuition will help in the application of DPs to mixture models. Stick-breaking Construction We have already intuited that draws from a DP are composed of a weighted sum of point masses. [8] made this precise by providing a constructive definition of the DP as such, called the stick-breaking construction. This construction is also significantly more straightforward and general than previous proofs of the existence of DPs. It is simply given as follows: β k Beta(1, α) θk H k 1 π k = β k (1 β l ) G = π k δ θ k (13) l=1 Then G DP(α, H). The construction of π can be understood metaphorically as follows. Starting with a stick of length 1, we break it at β 1, assigning π 1 to be the length of stick we just broke off. Now recursively break the other portion to obtain π 2, π 3 and so forth. The stick-breaking distribution over π is sometimes written π GEM(α), where the letters stand for Griffiths, Engen and McCloskey [10]. Because of its simplicity, the stick-breaking construction has lead to a variety of extensions as well as novel inference techniques for the Dirichlet process [11]. k=1 8

9 Applications Because of its simplicity, DPs are used across a wide variety of applications of Bayesian analysis in both statistics and machine learning. The simplest and most prevalent applications include: Bayesian model validation, density estimation and clustering via mixture models. We shall briefly describe the first two classes before detailing DP mixture models. How does one validate that a model gives a good fit to some observed data? The Bayesian approach would usually involve computing the marginal probability of the observed data under the model, and comparing this marginal probability to that for other models. If the marginal probability of the model of interest is highest we may conclude that we have a good fit. The choice of models to compare against is an issue in this approach, since it is desirable to compare against as large a class of models as possible. The Bayesian nonparametric approach gives an answer to this question: use the space of all possible distributions as our comparison class, with a prior over distributions. The DP is a popular choice for this prior, due to its simplicity, wide coverage of the class of all distributions, and recent advances in computationally efficient inference in DP models. The approach is usually to use the given parametric model as the base distribution of the DP, with the DP serving as a nonparametric relaxation around this parametric model. If the parametric model performs as well or better than the DP relaxed model, we have convincing evidence of the validity of the model. Another application of DPs is in density estimation [12, 5, 3, 6]. Here we are interested in modeling the density from which a given set of observations is drawn. To avoid limiting ourselves to any parametric class, we may again use a nonparametric prior over all densities. Here again DPs are a popular. However note that distributions drawn from a DP are discrete, thus do not have densities. The solution is to smooth out draws from the DP with a kernel. Let G DP(α, H) and let f(x θ) be a family of densities (kernels) indexed by θ. We use the following as our nonparametric density of x: p(x) = f(x θ)g(θ) dθ (14) Similarly, smoothing out DPs in this way is also useful in the nonparametric relaxation setting above. As we see below, this way of smoothing out DPs is equivalent to DP mixture models, if the data distributions F (θ) below are smooth with densities given by f(x θ). Dirichlet Process Mixture Models The most common application of the Dirichlet process is in clustering data using mixture models [12, 5, 3, 6]. Here the nonparametric nature of the Dirichlet process translates to mixture models with a countably infinite number of components. We model a set of observations {x 1,..., x n } using a set of latent parameters {θ 1,..., θ n }. Each θ i is drawn independently and identically from 9

10 G, while each x i has distribution F (θ i ) parametrized by θ i : x i θ i F (θ i ) θ i G G G α, H DP(α, H) (15) Because G is discrete, multiple θ i s can take on the same value simultaneously, and the above model can be seen as a mixture model, where x i s with the same value of θ i belong to the same cluster. The mixture perspective can be made more in agreement with the usual representation of mixture models using the stick-breaking construction (13). Let z i be a cluster assignment variable, which takes on value k with probability π k. Then (15) can be equivalently expressed as π α GEM(α) θ k H H z i π Mult(π) x i z i, {θ k} F (θ z i ) (16) with G = k=1 π kδ θ k and θ i = θz i. In mixture modeling terminology, π is the mixing proportion, θk are the cluster parameters, F (θ k ) is the distribution over data in cluster k, and H the prior over cluster parameters. The DP mixture model is an infinite mixture model a mixture model with a countably infinite number of clusters. However, because the π k s decrease exponentially quickly, only a small number of clusters will be used to model the data a priori (in fact, as we saw previously, the expected number of components used a priori is logarithmic in the number of observations). This is different than a finite mixture model, which uses a fixed number of clusters to model the data. In the DP mixture model, the actual number of clusters used to model data is not fixed, and can be automatically inferred from data using the usual Bayesian posterior inference framework (see [4] for a survey of MCMC inference procedures for DP mixture models). The equivalent operation for finite mixture models would be model averaging or model selection for the appropriate number of components, an approach which is fraught with difficulties. Thus infinite mixture models as exemplified by DP mixture models provide a compelling alternative to the traditional finite mixture model paradigm. Generalizations and Extensions The DP is the canonical distribution over probability measures and a wide range of generalizations have been proposed in the literature. First and foremost is the Pitman-Yor process [13, 11], which has recently seen successful applications modeling data exhibiting power-law properties [14, 15]. The Pitman-Yor process includes a third parameter d [0, 1), with d = 0 reducing to the DP. The various representations of the DP, including the Chinese restaurant process and the stick-breaking construction, have analogues for the Pitman-Yor process. Other generalizations of the DP are obtained by generalizing one of 10

11 its representations. These include Pólya trees, normalized random measure, Poisson-Kingman models, species sampling models and stick-breaking priors. The DP has also been used in more complex models involving more than one random probability measure. For example, in nonparametric regression we might have one probability measure for each value of a covariate, and in multitask settings each task might be associated with a probability measure with dependence across tasks implemented using a hierarchical Bayesian model. In the first situation the class of models is typically called dependent Dirichlet processes [16], while in the second the appropriate model is a hierarchical Dirichlet process [17]. Future Directions The Dirichlet process, and Bayesian nonparametrics in general, is an active area of research within both machine learning and statistics. Current research trends span a number of directions. Firstly there is the issue of efficient inference in DP models. [4] is an excellent survey of the state-of-the-art in 2000, with all algorithms based on Gibbs sampling or small-step Metropolis-Hastings MCMC sampling. Since then there has been much work, including split-and-merge and large-step auxiliary variable MCMC sampling, sequential Monte Carlo, expectation propagation, and variational methods. Secondly there has been interest in extending the DP, both in terms of new random distributions, as well as novel classes of nonparametric objects inspired by the DP. Thirdly, theoretical issues of convergence and consistency are being explored to provide frequentist guarantees for Bayesian nonparametric models. Finally there are applications of such models, to clustering, transfer learning, relational learning, models of cognition, sequence learning, and regression and classification among others. We believe DPs and Bayesian nonparametrics will prove to be rich and fertile grounds for research for years to come. Cross References Bayesian Methods, Prior Probabilities, Bayesian Nonparametrics. Recommended Reading In addition to the references embedded in the text above, we recommend the book [18] on Bayesian nonparametrics. [1] T. S. Ferguson. A Bayesian analysis of some nonparametric problems. Annals of Statistics, 1(2): , [2] W. J. Ewens. The sampling theory of selectively neutral alleles. Theoretical Population Biology, 3:87 112,

12 [3] M. D. Escobar and M. West. Bayesian density estimation and inference using mixtures. Journal of the American Statistical Association, 90: , [4] R. M. Neal. Markov chain sampling methods for Dirichlet process mixture models. Journal of Computational and Graphical Statistics, 9: , [5] R. M. Neal. Bayesian mixture modeling. In Proceedings of the Workshop on Maximum Entropy and Bayesian Methods of Statistical Analysis, volume 11, pages , [6] C. E. Rasmussen. The infinite Gaussian mixture model. In Advances in Neural Information Processing Systems, volume 12, [7] D. Blackwell and J. B. MacQueen. Ferguson distributions via Pólya urn schemes. Annals of Statistics, 1: , [8] J. Sethuraman. A constructive definition of Dirichlet priors. Statistica Sinica, 4: , [9] D. Aldous. Exchangeability and related topics. In École d Été de Probabilités de Saint-Flour XIII 1983, pages Springer, Berlin, [10] J. Pitman. Combinatorial stochastic processes. Technical Report 621, Department of Statistics, University of California at Berkeley, Lecture notes for St. Flour Summer School. [11] H. Ishwaran and L. F. James. Gibbs sampling methods for stick-breaking priors. Journal of the American Statistical Association, 96(453): , [12] A.Y. Lo. On a class of bayesian nonparametric estimates: I. density estimates. Annals of Statistics, 12(1): , [13] J. Pitman and M. Yor. The two-parameter Poisson-Dirichlet distribution derived from a stable subordinator. Annals of Probability, 25: , [14] S. Goldwater, T.L. Griffiths, and M. Johnson. Interpolating between types and tokens by estimating power-law generators. In Advances in Neural Information Processing Systems, volume 18, [15] Y. W. Teh. A hierarchical Bayesian language model based on Pitman-Yor processes. In Proceedings of the 21st International Conference on Computational Linguistics and 44th Annual Meeting of the Association for Computational Linguistics, pages , [16] S. MacEachern. Dependent nonparametric processes. In Proceedings of the Section on Bayesian Statistical Science. American Statistical Association,

13 [17] Y. W. Teh, M. I. Jordan, M. J. Beal, and D. M. Blei. Hierarchical Dirichlet processes. Journal of the American Statistical Association, 101(476): , [18] N. Hjort, C. Holmes, P. Müller, and S. Walker, editors. Bayesian Nonparametrics. Number 28 in Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press,

### Dirichlet Processes A gentle tutorial

Dirichlet Processes A gentle tutorial SELECT Lab Meeting October 14, 2008 Khalid El-Arini Motivation We are given a data set, and are told that it was generated from a mixture of Gaussian distributions.

### Bayesian Statistics: Indian Buffet Process

Bayesian Statistics: Indian Buffet Process Ilker Yildirim Department of Brain and Cognitive Sciences University of Rochester Rochester, NY 14627 August 2012 Reference: Most of the material in this note

### HT2015: SC4 Statistical Data Mining and Machine Learning

HT2015: SC4 Statistical Data Mining and Machine Learning Dino Sejdinovic Department of Statistics Oxford http://www.stats.ox.ac.uk/~sejdinov/sdmml.html Bayesian Nonparametrics Parametric vs Nonparametric

### The Chinese Restaurant Process

COS 597C: Bayesian nonparametrics Lecturer: David Blei Lecture # 1 Scribes: Peter Frazier, Indraneel Mukherjee September 21, 2007 In this first lecture, we begin by introducing the Chinese Restaurant Process.

### Dirichlet Process Gaussian Mixture Models: Choice of the Base Distribution

Görür D, Rasmussen CE. Dirichlet process Gaussian mixture models: Choice of the base distribution. JOURNAL OF COMPUTER SCIENCE AND TECHNOLOGY 5(4): 615 66 July 010/DOI 10.1007/s11390-010-1051-1 Dirichlet

### An Alternative Prior Process for Nonparametric Bayesian Clustering

An Alternative Prior Process for Nonparametric Bayesian Clustering Hanna M. Wallach Department of Computer Science University of Massachusetts Amherst Shane T. Jensen Department of Statistics The Wharton

### Lecture 6: The Bayesian Approach

Lecture 6: The Bayesian Approach What Did We Do Up to Now? We are given a model Log-linear model, Markov network, Bayesian network, etc. This model induces a distribution P(X) Learning: estimate a set

### Nonparametric Factor Analysis with Beta Process Priors

Nonparametric Factor Analysis with Beta Process Priors John Paisley Lawrence Carin Department of Electrical & Computer Engineering Duke University, Durham, NC 7708 jwp4@ee.duke.edu lcarin@ee.duke.edu Abstract

### STA 4273H: Statistical Machine Learning

STA 4273H: Statistical Machine Learning Russ Salakhutdinov Department of Statistics! rsalakhu@utstat.toronto.edu! http://www.cs.toronto.edu/~rsalakhu/ Lecture 6 Three Approaches to Classification Construct

### Non-parametric Bayesian Methods

Non-parametric Bayesian Methods Uncertainty in Artificial Intelligence Tutorial July 25 Zoubin Ghahramani Gatsby Computational Neuroscience Unit 1 University College London, UK Center for Automated Learning

### Basics of Statistical Machine Learning

CS761 Spring 2013 Advanced Machine Learning Basics of Statistical Machine Learning Lecturer: Xiaojin Zhu jerryzhu@cs.wisc.edu Modern machine learning is rooted in statistics. You will find many familiar

### Microclustering: When the Cluster Sizes Grow Sublinearly with the Size of the Data Set

Microclustering: When the Cluster Sizes Grow Sublinearly with the Size of the Data Set Jeffrey W. Miller Brenda Betancourt Abbas Zaidi Hanna Wallach Rebecca C. Steorts Abstract Most generative models for

Statistics Graduate Courses STAT 7002--Topics in Statistics-Biological/Physical/Mathematics (cr.arr.).organized study of selected topics. Subjects and earnable credit may vary from semester to semester.

### The Exponential Family

The Exponential Family David M. Blei Columbia University November 3, 2015 Definition A probability density in the exponential family has this form where p.x j / D h.x/ expf > t.x/ a./g; (1) is the natural

### Stick-breaking Construction for the Indian Buffet Process

Stick-breaking Construction for the Indian Buffet Process Yee Whye Teh Gatsby Computational Neuroscience Unit University College London 17 Queen Square London WC1N 3AR, UK ywteh@gatsby.ucl.ac.uk Dilan

### Machine Learning and Data Analysis overview. Department of Cybernetics, Czech Technical University in Prague. http://ida.felk.cvut.

Machine Learning and Data Analysis overview Jiří Kléma Department of Cybernetics, Czech Technical University in Prague http://ida.felk.cvut.cz psyllabus Lecture Lecturer Content 1. J. Kléma Introduction,

### The Basics of Graphical Models

The Basics of Graphical Models David M. Blei Columbia University October 3, 2015 Introduction These notes follow Chapter 2 of An Introduction to Probabilistic Graphical Models by Michael Jordan. Many figures

### Thomas L. Griffiths. Department of Psychology. University of California, Berkeley. Adam N. Sanborn. Department of Psychological and Brain Sciences

Categorization as nonparametric Bayesian density estimation 1 Running head: CATEGORIZATION AS NONPARAMETRIC BAYESIAN DENSITY ESTIMATION Categorization as nonparametric Bayesian density estimation Thomas

### Modeling Transfer Learning in Human Categorization with the Hierarchical Dirichlet Process

Modeling Transfer Learning in Human Categorization with the Hierarchical Dirichlet Process Kevin R. Canini kevin@cs.berkeley.edu Computer Science Division, University of California, Berkeley, CA 94720

### Tutorial on Markov Chain Monte Carlo

Tutorial on Markov Chain Monte Carlo Kenneth M. Hanson Los Alamos National Laboratory Presented at the 29 th International Workshop on Bayesian Inference and Maximum Entropy Methods in Science and Technology,

### Probabilistic Models for Big Data. Alex Davies and Roger Frigola University of Cambridge 13th February 2014

Probabilistic Models for Big Data Alex Davies and Roger Frigola University of Cambridge 13th February 2014 The State of Big Data Why probabilistic models for Big Data? 1. If you don t have to worry about

### CHAPTER 2 Estimating Probabilities

CHAPTER 2 Estimating Probabilities Machine Learning Copyright c 2016. Tom M. Mitchell. All rights reserved. *DRAFT OF January 24, 2016* *PLEASE DO NOT DISTRIBUTE WITHOUT AUTHOR S PERMISSION* This is a

### Topic models for Sentiment analysis: A Literature Survey

Topic models for Sentiment analysis: A Literature Survey Nikhilkumar Jadhav 123050033 June 26, 2014 In this report, we present the work done so far in the field of sentiment analysis using topic models.

### EC 6310: Advanced Econometric Theory

EC 6310: Advanced Econometric Theory July 2008 Slides for Lecture on Bayesian Computation in the Nonlinear Regression Model Gary Koop, University of Strathclyde 1 Summary Readings: Chapter 5 of textbook.

### Bayesian Models of Graphs, Arrays and Other Exchangeable Random Structures

Bayesian Models of Graphs, Arrays and Other Exchangeable Random Structures Peter Orbanz and Daniel M. Roy Abstract. The natural habitat of most Bayesian methods is data represented by exchangeable sequences

### Bootstrapping Big Data

Bootstrapping Big Data Ariel Kleiner Ameet Talwalkar Purnamrita Sarkar Michael I. Jordan Computer Science Division University of California, Berkeley {akleiner, ameet, psarkar, jordan}@eecs.berkeley.edu

### A Probabilistic Model for Online Document Clustering with Application to Novelty Detection

A Probabilistic Model for Online Document Clustering with Application to Novelty Detection Jian Zhang School of Computer Science Cargenie Mellon University Pittsburgh, PA 15213 jian.zhang@cs.cmu.edu Zoubin

### Information Theory and Coding Prof. S. N. Merchant Department of Electrical Engineering Indian Institute of Technology, Bombay

Information Theory and Coding Prof. S. N. Merchant Department of Electrical Engineering Indian Institute of Technology, Bombay Lecture - 17 Shannon-Fano-Elias Coding and Introduction to Arithmetic Coding

### Sampling via Moment Sharing: A New Framework for Distributed Bayesian Inference for Big Data

Sampling via Moment Sharing: A New Framework for Distributed Bayesian Inference for Big Data (Oxford) in collaboration with: Minjie Xu, Jun Zhu, Bo Zhang (Tsinghua) Balaji Lakshminarayanan (Gatsby) Bayesian

### Learning outcomes. Knowledge and understanding. Competence and skills

Syllabus Master s Programme in Statistics and Data Mining 120 ECTS Credits Aim The rapid growth of databases provides scientists and business people with vast new resources. This programme meets the challenges

### Gaussian Processes in Machine Learning

Gaussian Processes in Machine Learning Carl Edward Rasmussen Max Planck Institute for Biological Cybernetics, 72076 Tübingen, Germany carl@tuebingen.mpg.de WWW home page: http://www.tuebingen.mpg.de/ carl

### Parametric Models Part I: Maximum Likelihood and Bayesian Density Estimation

Parametric Models Part I: Maximum Likelihood and Bayesian Density Estimation Selim Aksoy Department of Computer Engineering Bilkent University saksoy@cs.bilkent.edu.tr CS 551, Fall 2015 CS 551, Fall 2015

### Statistical Machine Learning

Statistical Machine Learning UoC Stats 37700, Winter quarter Lecture 4: classical linear and quadratic discriminants. 1 / 25 Linear separation For two classes in R d : simple idea: separate the classes

### Handling attrition and non-response in longitudinal data

Longitudinal and Life Course Studies 2009 Volume 1 Issue 1 Pp 63-72 Handling attrition and non-response in longitudinal data Harvey Goldstein University of Bristol Correspondence. Professor H. Goldstein

### Service courses for graduate students in degree programs other than the MS or PhD programs in Biostatistics.

Course Catalog In order to be assured that all prerequisites are met, students must acquire a permission number from the education coordinator prior to enrolling in any Biostatistics course. Courses are

### Set theory, and set operations

1 Set theory, and set operations Sayan Mukherjee Motivation It goes without saying that a Bayesian statistician should know probability theory in depth to model. Measure theory is not needed unless we

### CS 2750 Machine Learning. Lecture 1. Machine Learning. http://www.cs.pitt.edu/~milos/courses/cs2750/ CS 2750 Machine Learning.

Lecture Machine Learning Milos Hauskrecht milos@cs.pitt.edu 539 Sennott Square, x5 http://www.cs.pitt.edu/~milos/courses/cs75/ Administration Instructor: Milos Hauskrecht milos@cs.pitt.edu 539 Sennott

### Statistical Machine Learning from Data

Samy Bengio Statistical Machine Learning from Data 1 Statistical Machine Learning from Data Gaussian Mixture Models Samy Bengio IDIAP Research Institute, Martigny, Switzerland, and Ecole Polytechnique

### PS 271B: Quantitative Methods II. Lecture Notes

PS 271B: Quantitative Methods II Lecture Notes Langche Zeng zeng@ucsd.edu The Empirical Research Process; Fundamental Methodological Issues 2 Theory; Data; Models/model selection; Estimation; Inference.

### BayesX - Software for Bayesian Inference in Structured Additive Regression

BayesX - Software for Bayesian Inference in Structured Additive Regression Thomas Kneib Faculty of Mathematics and Economics, University of Ulm Department of Statistics, Ludwig-Maximilians-University Munich

### Inference on Phase-type Models via MCMC

Inference on Phase-type Models via MCMC with application to networks of repairable redundant systems Louis JM Aslett and Simon P Wilson Trinity College Dublin 28 th June 202 Toy Example : Redundant Repairable

### Probabilities and Random Variables

Probabilities and Random Variables This is an elementary overview of the basic concepts of probability theory. 1 The Probability Space The purpose of probability theory is to model random experiments so

Graduate Programs in Statistics Course Titles STAT 100 CALCULUS AND MATR IX ALGEBRA FOR STATISTICS. Differential and integral calculus; infinite series; matrix algebra STAT 195 INTRODUCTION TO MATHEMATICAL

### Multivariate Normal Distribution

Multivariate Normal Distribution Lecture 4 July 21, 2011 Advanced Multivariate Statistical Methods ICPSR Summer Session #2 Lecture #4-7/21/2011 Slide 1 of 41 Last Time Matrices and vectors Eigenvalues

### Nonparametric adaptive age replacement with a one-cycle criterion

Nonparametric adaptive age replacement with a one-cycle criterion P. Coolen-Schrijner, F.P.A. Coolen Department of Mathematical Sciences University of Durham, Durham, DH1 3LE, UK e-mail: Pauline.Schrijner@durham.ac.uk

### Linear Threshold Units

Linear Threshold Units w x hx (... w n x n w We assume that each feature x j and each weight w j is a real number (we will relax this later) We will study three different algorithms for learning linear

### Selection Sampling from Large Data sets for Targeted Inference in Mixture Modeling

Selection Sampling from Large Data sets for Targeted Inference in Mixture Modeling Ioanna Manolopoulou, Cliburn Chan and Mike West December 29, 2009 Abstract One of the challenges of Markov chain Monte

### Spatial Statistics Chapter 3 Basics of areal data and areal data modeling

Spatial Statistics Chapter 3 Basics of areal data and areal data modeling Recall areal data also known as lattice data are data Y (s), s D where D is a discrete index set. This usually corresponds to data

### LÉVY-DRIVEN PROCESSES IN BAYESIAN NONPARAMETRIC INFERENCE

Bol. Soc. Mat. Mexicana (3) Vol. 19, 2013 LÉVY-DRIVEN PROCESSES IN BAYESIAN NONPARAMETRIC INFERENCE LUIS E. NIETO-BARAJAS ABSTRACT. In this article we highlight the important role that Lévy processes have

### Logistic Regression (1/24/13)

STA63/CBB540: Statistical methods in computational biology Logistic Regression (/24/3) Lecturer: Barbara Engelhardt Scribe: Dinesh Manandhar Introduction Logistic regression is model for regression used

### Gaussian Processes to Speed up Hamiltonian Monte Carlo

Gaussian Processes to Speed up Hamiltonian Monte Carlo Matthieu Lê Murray, Iain http://videolectures.net/mlss09uk_murray_mcmc/ Rasmussen, Carl Edward. "Gaussian processes to speed up hybrid Monte Carlo

### Producing Power-Law Distributions and Damping Word Frequencies with Two-Stage Language Models

Journal of Machine Learning Research 12 (2011) 2335-2382 Submitted 1/10; Revised 8/10; Published 7/11 Producing Power-Law Distributions and Damping Word Frequencies with Two-Stage Language Models Sharon

### arxiv:1112.0829v1 [math.pr] 5 Dec 2011

How Not to Win a Million Dollars: A Counterexample to a Conjecture of L. Breiman Thomas P. Hayes arxiv:1112.0829v1 [math.pr] 5 Dec 2011 Abstract Consider a gambling game in which we are allowed to repeatedly

### Bayesian Image Super-Resolution

Bayesian Image Super-Resolution Michael E. Tipping and Christopher M. Bishop Microsoft Research, Cambridge, U.K..................................................................... Published as: Bayesian

### Model-based Synthesis. Tony O Hagan

Model-based Synthesis Tony O Hagan Stochastic models Synthesising evidence through a statistical model 2 Evidence Synthesis (Session 3), Helsinki, 28/10/11 Graphical modelling The kinds of models that

### MASSACHUSETTS INSTITUTE OF TECHNOLOGY 6.436J/15.085J Fall 2008 Lecture 5 9/17/2008 RANDOM VARIABLES

MASSACHUSETTS INSTITUTE OF TECHNOLOGY 6.436J/15.085J Fall 2008 Lecture 5 9/17/2008 RANDOM VARIABLES Contents 1. Random variables and measurable functions 2. Cumulative distribution functions 3. Discrete

### Dynamic Linear Models with R

Giovanni Petris, Sonia Petrone, Patrizia Campagnoli Dynamic Linear Models with R SPIN Springer s internal project number, if known Monograph August 10, 2007 Springer Berlin Heidelberg NewYork Hong Kong

### An Introduction to Using WinBUGS for Cost-Effectiveness Analyses in Health Economics

Slide 1 An Introduction to Using WinBUGS for Cost-Effectiveness Analyses in Health Economics Dr. Christian Asseburg Centre for Health Economics Part 1 Slide 2 Talk overview Foundations of Bayesian statistics

### Learning outcomes. Knowledge and understanding. Ability and Competences. Evaluation capability and scientific approach

Syllabus Master s Programme in Statistics and Data Mining 120 ECTS Credits Aim The rapid growth of databases provides scientists and business people with vast new resources. This programme meets the challenges

### 1 Short Introduction to Time Series

ECONOMICS 7344, Spring 202 Bent E. Sørensen January 24, 202 Short Introduction to Time Series A time series is a collection of stochastic variables x,.., x t,.., x T indexed by an integer value t. The

### General theory of stochastic processes

CHAPTER 1 General theory of stochastic processes 1.1. Definition of stochastic process First let us recall the definition of a random variable. A random variable is a random number appearing as a result

### 6.4 The Infinitely Many Alleles Model

6.4. THE INFINITELY MANY ALLELES MODEL 93 NE µ [F (X N 1 ) F (µ)] = + 1 i

### Using Mixtures-of-Distributions models to inform farm size selection decisions in representative farm modelling. Philip Kostov and Seamus McErlean

Using Mixtures-of-Distributions models to inform farm size selection decisions in representative farm modelling. by Philip Kostov and Seamus McErlean Working Paper, Agricultural and Food Economics, Queen

### I, ONLINE KNOWLEDGE TRANSFER FREELANCER PORTAL M. & A.

ONLINE KNOWLEDGE TRANSFER FREELANCER PORTAL M. Micheal Fernandas* & A. Anitha** * PG Scholar, Department of Master of Computer Applications, Dhanalakshmi Srinivasan Engineering College, Perambalur, Tamilnadu

### Anomaly detection for Big Data, networks and cyber-security

Anomaly detection for Big Data, networks and cyber-security Patrick Rubin-Delanchy University of Bristol & Heilbronn Institute for Mathematical Research Joint work with Nick Heard (Imperial College London),

### Master s Theory Exam Spring 2006

Spring 2006 This exam contains 7 questions. You should attempt them all. Each question is divided into parts to help lead you through the material. You should attempt to complete as much of each problem

### 11. Time series and dynamic linear models

11. Time series and dynamic linear models Objective To introduce the Bayesian approach to the modeling and forecasting of time series. Recommended reading West, M. and Harrison, J. (1997). models, (2 nd

### Computing with Finite and Infinite Networks

Computing with Finite and Infinite Networks Ole Winther Theoretical Physics, Lund University Sölvegatan 14 A, S-223 62 Lund, Sweden winther@nimis.thep.lu.se Abstract Using statistical mechanics results,

### Estimating Incentive and Selection Effects in Medigap Insurance. Market: An Application with Dirichlet Process Mixture Model

Estimating Incentive and Selection Effects in Medigap Insurance Market: An Application with Dirichlet Process Mixture Model Xuequn Hu Department of Economics University of South Florida Email: xhu8@usf.edu

### Monte Carlo and Empirical Methods for Stochastic Inference (MASM11/FMS091)

Monte Carlo and Empirical Methods for Stochastic Inference (MASM11/FMS091) Magnus Wiktorsson Centre for Mathematical Sciences Lund University, Sweden Lecture 6 Sequential Monte Carlo methods II February

### Square Root Propagation

Square Root Propagation Andrew G. Howard Department of Computer Science Columbia University New York, NY 10027 ahoward@cs.columbia.edu Tony Jebara Department of Computer Science Columbia University New

### CONTENTS OF DAY 2. II. Why Random Sampling is Important 9 A myth, an urban legend, and the real reason NOTES FOR SUMMER STATISTICS INSTITUTE COURSE

1 2 CONTENTS OF DAY 2 I. More Precise Definition of Simple Random Sample 3 Connection with independent random variables 3 Problems with small populations 8 II. Why Random Sampling is Important 9 A myth,

### Data Mining Chapter 6: Models and Patterns Fall 2011 Ming Li Department of Computer Science and Technology Nanjing University

Data Mining Chapter 6: Models and Patterns Fall 2011 Ming Li Department of Computer Science and Technology Nanjing University Models vs. Patterns Models A model is a high level, global description of a

### Machine Learning and Pattern Recognition Logistic Regression

Machine Learning and Pattern Recognition Logistic Regression Course Lecturer:Amos J Storkey Institute for Adaptive and Neural Computation School of Informatics University of Edinburgh Crichton Street,

### The Dirichlet-Multinomial and Dirichlet-Categorical models for Bayesian inference

The Dirichlet-Multinomial and Dirichlet-Categorical models for Bayesian inference Stephen Tu tu.stephenl@gmail.com 1 Introduction This document collects in one place various results for both the Dirichlet-multinomial

### Introduction to Markov Chain Monte Carlo

Introduction to Markov Chain Monte Carlo Monte Carlo: sample from a distribution to estimate the distribution to compute max, mean Markov Chain Monte Carlo: sampling using local information Generic problem

### Introduction to latent variable models

Introduction to latent variable models lecture 1 Francesco Bartolucci Department of Economics, Finance and Statistics University of Perugia, IT bart@stat.unipg.it Outline [2/24] Latent variables and their

### Sample Size Calculation for Finding Unseen Species

Bayesian Analysis (2009) 4, Number 4, pp. 763 792 Sample Size Calculation for Finding Unseen Species Hongmei Zhang and Hal Stern Abstract. Estimation of the number of species extant in a geographic region

### 1 Prior Probability and Posterior Probability

Math 541: Statistical Theory II Bayesian Approach to Parameter Estimation Lecturer: Songfeng Zheng 1 Prior Probability and Posterior Probability Consider now a problem of statistical inference in which

### Least Squares Estimation

Least Squares Estimation SARA A VAN DE GEER Volume 2, pp 1041 1045 in Encyclopedia of Statistics in Behavioral Science ISBN-13: 978-0-470-86080-9 ISBN-10: 0-470-86080-4 Editors Brian S Everitt & David

### 1 Limiting distribution for a Markov chain

Copyright c 2009 by Karl Sigman Limiting distribution for a Markov chain In these Lecture Notes, we shall study the limiting behavior of Markov chains as time n In particular, under suitable easy-to-check

### 3. Examples of discrete probability spaces.. Here α ( j)

3. EXAMPLES OF DISCRETE PROBABILITY SPACES 13 3. Examples of discrete probability spaces Example 17. Toss n coins. We saw this before, but assumed that the coins are fair. Now we do not. The sample space

### Markov Chain Monte Carlo Simulation Made Simple

Markov Chain Monte Carlo Simulation Made Simple Alastair Smith Department of Politics New York University April2,2003 1 Markov Chain Monte Carlo (MCMC) simualtion is a powerful technique to perform numerical

### Nonparametric Bayesian Co-clustering Ensembles

Nonparametric Bayesian Co-clustering Ensembles Pu Wang Kathryn B. Lasey Carlotta Domeniconi Michael I. Jordan Abstract A nonparametric Bayesian approach to co-clustering ensembles is presented. Similar

### Streaming Variational Inference for Bayesian Nonparametric Mixture Models

Streaming Variational Inference for Bayesian Nonparametric Mixture Models Alex Tank Nicholas J. Foti Emily B. Fox University of Washington University of Washington University of Washington Department of

### Ideal Class Group and Units

Chapter 4 Ideal Class Group and Units We are now interested in understanding two aspects of ring of integers of number fields: how principal they are (that is, what is the proportion of principal ideals

### U.C. Berkeley CS276: Cryptography Handout 0.1 Luca Trevisan January, 2009. Notes on Algebra

U.C. Berkeley CS276: Cryptography Handout 0.1 Luca Trevisan January, 2009 Notes on Algebra These notes contain as little theory as possible, and most results are stated without proof. Any introductory

### The Optimality of Naive Bayes

The Optimality of Naive Bayes Harry Zhang Faculty of Computer Science University of New Brunswick Fredericton, New Brunswick, Canada email: hzhang@unbca E3B 5A3 Abstract Naive Bayes is one of the most

### 4. Factor polynomials over complex numbers, describe geometrically, and apply to real-world situations. 5. Determine and apply relationships among syn

I The Real and Complex Number Systems 1. Identify subsets of complex numbers, and compare their structural characteristics. 2. Compare and contrast the properties of real numbers with the properties of

### WHAT ARE MATHEMATICAL PROOFS AND WHY THEY ARE IMPORTANT?

WHAT ARE MATHEMATICAL PROOFS AND WHY THEY ARE IMPORTANT? introduction Many students seem to have trouble with the notion of a mathematical proof. People that come to a course like Math 216, who certainly

Adaptive Online Gradient Descent Peter L Bartlett Division of Computer Science Department of Statistics UC Berkeley Berkeley, CA 94709 bartlett@csberkeleyedu Elad Hazan IBM Almaden Research Center 650

### Modern Optimization Methods for Big Data Problems MATH11146 The University of Edinburgh

Modern Optimization Methods for Big Data Problems MATH11146 The University of Edinburgh Peter Richtárik Week 3 Randomized Coordinate Descent With Arbitrary Sampling January 27, 2016 1 / 30 The Problem

### Monte Carlo and Empirical Methods for Stochastic Inference (MASM11/FMS091)

Monte Carlo and Empirical Methods for Stochastic Inference (MASM11/FMS091) Magnus Wiktorsson Centre for Mathematical Sciences Lund University, Sweden Lecture 5 Sequential Monte Carlo methods I February

### Prime Numbers. Chapter Primes and Composites

Chapter 2 Prime Numbers The term factoring or factorization refers to the process of expressing an integer as the product of two or more integers in a nontrivial way, e.g., 42 = 6 7. Prime numbers are

### An Introduction to Data Mining

An Introduction to Intel Beijing wei.heng@intel.com January 17, 2014 Outline 1 DW Overview What is Notable Application of Conference, Software and Applications Major Process in 2 Major Tasks in Detail

### CS 2750 Machine Learning. Lecture 1. Machine Learning. CS 2750 Machine Learning.

Lecture 1 Machine Learning Milos Hauskrecht milos@cs.pitt.edu 539 Sennott Square, x-5 http://www.cs.pitt.edu/~milos/courses/cs75/ Administration Instructor: Milos Hauskrecht milos@cs.pitt.edu 539 Sennott

### Course: Model, Learning, and Inference: Lecture 5

Course: Model, Learning, and Inference: Lecture 5 Alan Yuille Department of Statistics, UCLA Los Angeles, CA 90095 yuille@stat.ucla.edu Abstract Probability distributions on structured representation.

### A Learning Based Method for Super-Resolution of Low Resolution Images

A Learning Based Method for Super-Resolution of Low Resolution Images Emre Ugur June 1, 2004 emre.ugur@ceng.metu.edu.tr Abstract The main objective of this project is the study of a learning based method