top of page

Blog

Corporate Training, DEI Outcomes, and the Statistical Model We Often Ignore

Think of a corporate system through the lens of a statistical model. Conceptually, employee behavior and organizational outcomes can be represented as: Observed behavior/outcome = System effects + Individual characteristics + Random variation System effects include leadership, incentives, organizational structure, policies, hiring and promotion processes, accountability, and culture, etc. Individual characteristics include skills, experience, motivation, preferences, personal

The Propensity Score Controversy - Simulations

A few months ago, I wrote a blog post about the controversial use of propensity scores (PS). I believe it would be more informative to illustrate the key issues with supporting evidence from simulations. Let us consider a simple scenario in which the response of interest, π‘Œ, is a continuous variable associated with three prognostic covariates, 𝛸₁, 𝛸₂, and π‘ˆ. We assume that 𝛸₁ and 𝛸₂ are continuous and observed, whereas π‘ˆ is binary and unobserved. Suppose the outcome i

The Propensity Score Controversy

Propensity score (PS) methods are probably the most widely used statistical tools for causal inference in observational studies. In medical research, epidemiology, economics, political science, and, increasingly, data science, they are often perceived as a principled way to adjust for confounding when randomization is not available. Despite their popularity, propensity scores remain deeply controversial among statisticians. Some view them as an elegant design-based framework

From Likelihood to Loss: Why Statistics and Data Science Speak Different Languages

One persistent source of confusion for practitioners moving between statistics and data science is language. The two fields often describe the same ideas using different terms: covariates become features, parameters become weights, estimation becomes training, and likelihood becomes loss. At first glance, this may look like a fundamental divide. It is not. Much of the underlying mathematics is the same. The difference is largely a matter of objectives, traditions, and audienc

Randomization Is Not Just About Balance

Randomization in clinical trials is often perceived as a tool to β€œbalance covariates” between treatment groups. While this view is correct, it is incomplete and somewhat misleading. Randomization is not primarily about balance, but rather, it provides a design-based foundation for valid causal inference. Where Does the Probability Come From? In a randomized clinical trial (RCT), probability is not introduced through modeling assumptions such as normality, nor is it fundamenta

Misconceptions About Linear Regression Assumptions

I recently came across a LinkedIn post discussing the statistical assumptions of linear regression. Because the misconceptions in that post seem to be quite common, even among statisticians, I feel strongly compelled to write about them. The author claimed that the validity of linear regression depends on several key assumptions, namely: Linearity: The relationship between the dependent variable Y and the independent variable(s) X must be linear. Independence: The observation

A Bridge Between Regression and ANOVA Thinking

In dose-response studies, the dose level can be treated either as a classification variable in an ANOVA-type model or as a continuous variable in a regression model. There is a fun little bridge between these two seemingly different approaches, for example, when the underlying dose-response relationship can be represented using generalized linear models (GLMs). This post illustrates that bridge in the specific context of Gaussian linear models. It should be noted that scaling

What Defines a Good Clinical Statistician

I recently attended a leadership training session where a colleague shared an experience involving a physician on a Data Monitoring Committee (DMC) who questioned the qualifications of the committee statistician. The physician was concerned because the statistician often remained quiet during routine DMC meetings, which typically focus on safety issues. My colleague, however, was confident that the statistician was fully qualified for the role. This raises an interesting ques

The Conditional Error Principle and Adaptive Designs

The Conditional Error Principle (CEP) is the foundation of many frequentist adaptive designs. It asserts that any new statistical test chosen after an adaptation must have a conditional error rate no greater than that of the original, pre-specified test. This allows for flexible mid-course modifications to a trial, such as sample size increase, while still controlling the overall experimental-wise type I error rate. Formally, the CEP framework can be described as follows. Sta

Blinded Sample Size Re-estimation for Continuous Endpoints - Part 3

In Parts 1 and 2 of this series, we evaluated moment-based approaches for blinded sample size re-estimation (SSR). This part focuses on a likelihood-based alternative - the maximum likelihood estimation (MLE), which leverages the full data likelihood rather than relying solely on sample moments, such as variance or kurtosis. As before, a continuous observation 𝑋 from the pooled data of a two-group parallel study can be viewed as a random variable arising from a two-component

Andrew Yan

© 2026 by Andrew Yan

Powered and secured by Wix

Contact 

Ask me something

Thanks for submitting!

bottom of page