Some Contributions to Inference and Model/Variable Selection in High-Dimensional Problems
| dc.contributor.author | Paul, Sayantan | |
| dc.date.accessioned | 2026-09-18T06:29:20Z | |
| dc.date.issued | 2026-08-19 | |
| dc.description | This thesis has been completed under the supervision of Dr. Arijit Chakrabarti | |
| dc.description.abstract | In today’s world, we often come across situations where we need to infer about several parameters simultaneously based on a single dataset. This is known as the problem of simultaneous statistical inference and the need for such inference arises for data from the fields of biology, astronomy, genomics, bioinformatics, medicine, economics, finance, and image processing, just to name a few. As the name suggests, instead of being concerned about the optimality of inference for the individual parameters, the main interest here is in proposing procedures which can provide satisfactory results in the overall inferential problem involving all parameters of interest. The need for simultaneous inference may arise in the contexts of hypothesis testing, point estimation of several parameters, and building confidence intervals. Such inference becomes challenging when the number of parameters increases with sample size at the same or at a higher rate, i.e., when the problem becomes the high-dimensional. Equally challenging are the problems of model selection or variable selection for data of such kind. In this thesis, our interest is in simultaneous statistical inference and model/variable selection in the high-dimensional setting. One of our main goals would be to propose new inference and model/variable selection procedures and study their properties through theory and simulations. We would also employ some of our proposed methods in suitable real-life examples to understand how they perform in actual applications. We would also like to theoretically address interesting questions about existing procedures widely used in such contexts. An important focus of this work will be the special situations when the number of significant (which may mean nonzero or of sufficiently large magnitude in appropriate contexts) parameters is small compared to the total number of parameters under consideration. These are the so-called “sparse” situations. The assumption of sparsity is natural and very common in the literature, e.g., in the high-dimensional regression setting and in the problem of inference on a high dimensional mean vector. Our approaches here will be built under the Bayesian setting where the unknown parameters are assumed to follow some prior probability distributions. A big focus of our theoretical work will be on studying the optimality (in the frequentist or Bayesian decision theoretic sense) of the new methods proposed or of existing methods. A natural Bayesian approach under sparse settings is to model the parameters by two-group spikeand-slab priors. These priors are expressed as mixtures of a distribution degenerate at 0 (or highly concentrated near 0) and a distribution with high spread. However, as noted by many authors, inference with such priors can be computationally very challenging in high-dimensional situations and complex parametric frameworks. As an alternative to two-group mixture priors, there have been proposals in the literature to consider unimodal, continuous priors having sufficient mass around zero while possessing sufficiently heavy tails. Such priors are named as one-group “global-local” priors, the name “global-local” originating from the use of two sets of parameters to simultaneously enforce and modulate shrinkage at the global level and at the levels of individual parameters respectively. These are called the global shrinkage parameter and the local shrinkage parameters respectively. Although serious efforts have gone into studying optimality of inference using one-group priors in sparse parametric settings, several interesting questions in this connection are still unanswered and several areas somewhat less trodden in the current literature. This thesis is a modest attempt to address some of these in the context of specific models. A large part of our work will focus on inference on the famous normal means model and the linear regression model. The thesis concludes with a study of inference on non-normal count data. Our study is based on a broad class of one-group global-local shrinkage priors covering a lot of popularly used priors including the horseshoe. When the mean parameter of the sparse normal means model is modeled by a spike-and-slab prior, it has been shown in the literature that the corresponding posterior distribution contracts around the truth at a near minimax rate when the level of sparsity is unknown. This raises the question of whether the same phenomenon works when one uses a one-group prior instead of its two-group counterpart for the same problem. We provide an answer to this question in Chapter 2 of the thesis. In order to handle the unknown level of sparsity, we either estimate the global shrinkage parameter based on the data, an empirical Bayes approach, or model it by a non-degenerate absolutely continuous prior distribution on a suitable support in a full Bayes approach. We establish that the posterior distributions of the mean parameter vector, when modelled by broad classes of one-group priors, contract around the truth at a near minimax rate, for both the empirical Bayes and full Bayes approaches. Another interesting question is whether one-group priors can be used to form a good decision rule for the simultaneous testing problem of whether the means are zero or not when the means are truly generated from a two-group prior. In the latter half of this chapter, we are interested in answering this question, assuming that the level of sparsity is unknown. Considering a full Bayes approach, we are to establish that the Bayes risk of a decision rule using a broad class of one-group priors attains the Bayes risk of the optimal rule for the two-group settings asymptotically for a wide range of sparsity levels. The loss function considered is the additive symmetric 0 − 1 loss measuring the number of misclassifications made by a multiple testing rule. One of the main goals in a multiple hypothesis testing problem is to propose a decision rule which can control some overall measure of type I error rate, e.g., the False Discovery Rate (FDR). Given that a testing rule controls the FDR at some desired level, the next obvious question is whether the same rule can provide any control over the False Negative Rate (FNR). In this context, researchers are often interested in the optimal multiple testing rules in terms of controlling sum of certain type I and type II error measures. This can be answered by studying the minimax risk of multiple testing rules with respect to the corresponding loss functions. Very recently, for the normal means model, the expressions for the minimax risks based on misclassification (or Hamming) loss and the loss defined as the sum of FDP and FNP have been derived. It has also been proved that the famous Benjamini-Hochberg (BH) procedure and an ℓ−value based procedure using spike-and-slab priors attain the minimax risk asymptotically adaptively over broad sparsity levels. This motivates us to study whether decision rules based on one-group priors, if any, can enjoy such asymptotic optimality. When the level of sparsity is known, by choosing the global shrinkage parameter appropriately based on the knowledge of sparsity, we prove that the corresponding decision rule based on the broad class of one-group priors mentioned earlier achieves the minimax risk for both loss functions stated earlier. When the level of sparsity is unknown, some empirical Bayes and full Bayes versions of our decision rules can also attain the minimax risk. These results are proved in Chapter 3 of the thesis. Another important problem of interest is variable/model selection in a high-dimensional normal linear regression model. In this thesis, we are interested in a situation where covariates under study inregression model. In Chapter 4 of this thesis, motivated by the existing literature, we are interested in proposing a decision rule and resulting estimators, based on a general class of global-local priors, which have the “Oracle property”, in the sense that, they achieve variable selection consistency and optimal estimation rate, respectively. We propose a decision rule which declares a group to be active if the ratio of the ℓ2 norm of the posterior mean of the group regression coefficient to that of the least square estimate exceeds half. When the design matrix is block-orthogonal, we are able to establish that the global shrinkage parameter can be chosen in such a way that our proposed inference procedures have both selection consistency and optimal estimation rate, provided the level of sparsity is known. Even if the sparsity pattern is unknown, our modified decision rules, by either estimating the global shrinkage parameter from the data or by modeling a prior on it, can still enjoy the oracle property. In the simulation studies, our rules perform favorably compared to many existing methods in a variety of sparsity settings. Our methods, when applied to real datasets, also return encouraging results. In Chapters 2 through 4 of the thesis, our main areas of interest were the situations when the observed data are generated from normal distributions of various parametric forms. However, depending on the problem of interest, the Gaussianity assumption is not always proper. In Chapter 5 of this thesis, we are interested in one such example where the data consists of counts of events, most counts being close to zero while some are moderate or large. The data is modelled by Poisson distribution with unknown means. Clearly the natural prior for such means would be a two-group mixture with large mass on the component concentrated near zero. Our interest is in finding at if one-group modelling is still a good alternative in this context, as seen in previously in Chapters 2 through 4 of this thesis. Specifically, one of the questions this leads us to is whether any decision rule for multiple testing based on one-group priors can approximate the optimal rule with respect to two-group priors in terms of risk when the sample size gets large. Towards that we first obtain the asymptotic expression for the optimal Bayes risk under two-group prior in appropriate asymptotic framework. The loss is taken to be additive symmetric 0 − 1 loss. Next, irrespective of the sparsity pattern to be known or unknown, we establish that the Bayes risks corresponding to our proposed decision rules based on one-group priors attain the optimal Bayes risk, up to some multiplicative constant. Finally, the theoretical results are verified using simulation studies followed by a real data analysis. Many of the theoretical results derived in this thesis are the first of their kind in the literature in their specific contexts. Last but not the least, our theoretical results, simulations and to some extent real data analyses reinforce the logic of using appropriately chosen one-group priors as alternatives to their two-group counterparts in high-dimensional sparse parametric settings. | |
| dc.identifier.citation | 216p. | |
| dc.identifier.uri | http://hdl.handle.net/10263/7949 | |
| dc.language.iso | en | |
| dc.relation.ispartofseries | ISI Ph.D Thesis; TH697 | |
| dc.subject | Asymptotic minimaxity | |
| dc.subject | Near-minimax rate | |
| dc.subject | One-group prior | |
| dc.subject | Sparsity | |
| dc.subject | ABOS | |
| dc.subject | Half-thresholding | |
| dc.subject | Optimal estimation rate | |
| dc.subject | Variable selection consistency | |
| dc.subject | FDR | |
| dc.title | Some Contributions to Inference and Model/Variable Selection in High-Dimensional Problems | |
| dc.type | Thesis |
Files
License bundle
1 - 1 of 1
No Thumbnail Available
- Name:
- license.txt
- Size:
- 1.71 KB
- Format:
- Item-specific license agreed upon to submission
- Description:
