Increased Fisher’s information for parameters of association in count regression via extreme ranks

Daniel F. Linder; Jingjing Yin; Haresh Rochani; Hani Samawi; Sanjay Sethi

doi:10.1080/03610926.2017.1316859

Increased Fisher’s information for parameters of association in count regression via extreme ranks

Daniel F. Linder, Jingjing Yin, Haresh Rochani, Hani Samawi, Sanjay Sethi

Biostats & Data Science

Research output: Contribution to journal › Article › peer-review

1 Scopus citations

Abstract

The article details a sampling scheme which can lead to a reduction in sample size and cost in clinical and epidemiological studies of association between a count outcome and risk factor. We show that inference in two common generalized linear models for count data, Poisson and negative binomial regression, is improved by using a ranked auxiliary covariate, which guides the sampling procedure. This type of sampling has typically been used to improve inference on a population mean. The novelty of the current work is its extension to log-linear models and derivations showing that the sampling technique results in an increase in information as compared to simple random sampling. Specifically, we show that under the proposed sampling strategy the maximum likelihood estimate of the risk factor’s coefficient is improved through an increase in the Fisher’s information. A simulation study is performed to compare the mean squared error, bias, variance, and power of the sampling routine with simple random sampling under various data-generating scenarios. We also illustrate the merits of the sampling scheme on a real data set from a clinical setting of males with chronic obstructive pulmonary disease. Empirical results from the simulation study and data analysis coincide with the theoretical derivations, suggesting that a significant reduction in sample size, and hence study cost, can be realized while achieving the same precision as a simple random sample.

Original language	English (US)
Pages (from-to)	1181-1203
Number of pages	23
Journal	Communications in Statistics - Theory and Methods
Volume	47
Issue number	5
DOIs	https://doi.org/10.1080/03610926.2017.1316859
State	Published - Mar 4 2018

Keywords

Count regression
Fisher’s information
log-linear model
sample size
study cost

ASJC Scopus subject areas

Statistics and Probability

Access to Document

10.1080/03610926.2017.1316859

Cite this

@article{37fb12500f2341cfa27df90120170c31,

title = "Increased Fisher{\textquoteright}s information for parameters of association in count regression via extreme ranks",

abstract = "The article details a sampling scheme which can lead to a reduction in sample size and cost in clinical and epidemiological studies of association between a count outcome and risk factor. We show that inference in two common generalized linear models for count data, Poisson and negative binomial regression, is improved by using a ranked auxiliary covariate, which guides the sampling procedure. This type of sampling has typically been used to improve inference on a population mean. The novelty of the current work is its extension to log-linear models and derivations showing that the sampling technique results in an increase in information as compared to simple random sampling. Specifically, we show that under the proposed sampling strategy the maximum likelihood estimate of the risk factor{\textquoteright}s coefficient is improved through an increase in the Fisher{\textquoteright}s information. A simulation study is performed to compare the mean squared error, bias, variance, and power of the sampling routine with simple random sampling under various data-generating scenarios. We also illustrate the merits of the sampling scheme on a real data set from a clinical setting of males with chronic obstructive pulmonary disease. Empirical results from the simulation study and data analysis coincide with the theoretical derivations, suggesting that a significant reduction in sample size, and hence study cost, can be realized while achieving the same precision as a simple random sample.",

keywords = "Count regression, Fisher{\textquoteright}s information, log-linear model, sample size, study cost",

author = "Linder, {Daniel F.} and Jingjing Yin and Haresh Rochani and Hani Samawi and Sanjay Sethi",

note = "Publisher Copyright: {\textcopyright} 2018 Taylor & Francis Group, LLC.",

year = "2018",

month = mar,

day = "4",

doi = "10.1080/03610926.2017.1316859",

language = "English (US)",

volume = "47",

pages = "1181--1203",

journal = "Communications in Statistics - Theory and Methods",

issn = "0361-0926",

publisher = "Taylor and Francis Ltd.",

number = "5",

}

TY - JOUR

T1 - Increased Fisher’s information for parameters of association in count regression via extreme ranks

AU - Linder, Daniel F.

AU - Yin, Jingjing

AU - Rochani, Haresh

AU - Samawi, Hani

AU - Sethi, Sanjay

PY - 2018/3/4

Y1 - 2018/3/4

N2 - The article details a sampling scheme which can lead to a reduction in sample size and cost in clinical and epidemiological studies of association between a count outcome and risk factor. We show that inference in two common generalized linear models for count data, Poisson and negative binomial regression, is improved by using a ranked auxiliary covariate, which guides the sampling procedure. This type of sampling has typically been used to improve inference on a population mean. The novelty of the current work is its extension to log-linear models and derivations showing that the sampling technique results in an increase in information as compared to simple random sampling. Specifically, we show that under the proposed sampling strategy the maximum likelihood estimate of the risk factor’s coefficient is improved through an increase in the Fisher’s information. A simulation study is performed to compare the mean squared error, bias, variance, and power of the sampling routine with simple random sampling under various data-generating scenarios. We also illustrate the merits of the sampling scheme on a real data set from a clinical setting of males with chronic obstructive pulmonary disease. Empirical results from the simulation study and data analysis coincide with the theoretical derivations, suggesting that a significant reduction in sample size, and hence study cost, can be realized while achieving the same precision as a simple random sample.

AB - The article details a sampling scheme which can lead to a reduction in sample size and cost in clinical and epidemiological studies of association between a count outcome and risk factor. We show that inference in two common generalized linear models for count data, Poisson and negative binomial regression, is improved by using a ranked auxiliary covariate, which guides the sampling procedure. This type of sampling has typically been used to improve inference on a population mean. The novelty of the current work is its extension to log-linear models and derivations showing that the sampling technique results in an increase in information as compared to simple random sampling. Specifically, we show that under the proposed sampling strategy the maximum likelihood estimate of the risk factor’s coefficient is improved through an increase in the Fisher’s information. A simulation study is performed to compare the mean squared error, bias, variance, and power of the sampling routine with simple random sampling under various data-generating scenarios. We also illustrate the merits of the sampling scheme on a real data set from a clinical setting of males with chronic obstructive pulmonary disease. Empirical results from the simulation study and data analysis coincide with the theoretical derivations, suggesting that a significant reduction in sample size, and hence study cost, can be realized while achieving the same precision as a simple random sample.

KW - Count regression

KW - Fisher’s information

KW - log-linear model

KW - sample size

KW - study cost

UR - http://www.scopus.com/inward/record.url?scp=85029679634&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=85029679634&partnerID=8YFLogxK

U2 - 10.1080/03610926.2017.1316859

DO - 10.1080/03610926.2017.1316859

M3 - Article

AN - SCOPUS:85029679634

SN - 0361-0926

VL - 47

SP - 1181

EP - 1203

JO - Communications in Statistics - Theory and Methods

JF - Communications in Statistics - Theory and Methods

IS - 5

ER -

Increased Fisher’s information for parameters of association in count regression via extreme ranks

Abstract

Keywords

ASJC Scopus subject areas

Access to Document

Other files and links

Fingerprint

Cite this