Nonlinear Parameter Estimation, Confidence Region Minimization, and Optimization of Experimental Design
Chelise Van De Graaff and Larry Baxter
Department of Chemical Engineering
Brigham Young University
Linear regression is the traditional method of approximating the relationships that exist between a set of response and explanatory variables. However, the general linearized approach becomes faulty when 1) expressions are nonlinear in parameters, and 2) random errors are correlated with constant variance. Under these circumstances, nonlinear regression is an attractive alternative that provides more robust estimates of functional parameters and uncertainty.
When reporting the statistical significance of a model as predicted by data interpretation, it is common to report confidence intervals (parameter estimate +/- confidence interval) with standard error bars and an R2 value. This generalized approach provides some insight into the certainty of the governing model, but can be deceptive. Confidence intervals, by nature, do not include parameter correlation. It is more correct to report a statistically significant 'confidence region:' a contour that has the same number of dimensions as there are parameters. This confidence region is ellipsoidally shaped, and includes within its boundary a more robust estimate of joint parameter correlation.
If an a priori knowledge of the type of experimental correlation between the response and explanatory variables exists, traditional experimental design involves evenly spacing the location at which data will be collected. Spacing of collected data usually corresponds with the available number of data to be collected. For example, if an experiment can be replicated ten times producing ten unique data points, each data point is collected evenly spaced at one-tenth of the range over the entire experimental range. Heuristically this approach seems to most effectively utilize data collection, but in comparison, nonlinear analysis can predict with higher accuracy where data should be collected. Data should be collected based on the number of functional parameters, not the available number of data points. Efficient data collection is usually not linearly spaced over the available range, but instead collection falls in nonlinear patterns. For example, data collection for a cubic model usually is most efficient when clustered at one-third and two-thirds of the available range.
No comments:
Post a Comment