Skip to content

2022

T Statistics Program. It also presents the techniques used to pro

0925 Publ 16 (PDF) · 2026-10-03 edition · updated 2026-10-04 · United States

duce estimates of the total number of active corporations and any associated variables, as well as an assessment of the data limitations, including sampling and nonsampling errors.

Background

From TY 1916 through TY 1950, SOI extracted data from every corporate income tax return that was filed. In TY 1951, SOI introduced stratified probability sampling. Since then, the sample size has generally decreased while the corporate tax return population has increased. For example, for TY 1951, the sample accounted for 41.5% of the entire population, or 285,000 of the 687,000 total returns that were filed. For TY 2022, the sample accounted for about 1.83% of the total population of just over 7.5 million returns. This population count differs from the estimated population count cited elsewhere in this publication because the sampling frame includes out-ofscope and duplicate returns.

For TY 1951, SOI stratified the sample by size of total assets and industry. However, from TY 1952 through TY 1967, SOI stratified the sample by a measure of size only. The size

was measured by either business volume (TY 1953– TY 1958) or total assets (TY 1952 and TY 1959–TY1967). Since TY 1968, SOI has stratified returns by total assets, and for Forms 1120 and 1120-S by total assets and a measure of income. [1].

Target Population

The target population consists of all returns of active corporations organized for profit that are required to file one of the 1120-series forms included in this study.

Survey Population

The survey population includes corporate tax returns filed using one of the 1120-series forms selected for the study and posted to the IRS Business Master File (BMF). Amended returns and returns for which the tax liabilities changed because of a tax audit were excluded from the survey. Figure E gives the number of corporate returns by form type that were subject to sampling during TYs 2019 through 2022, as well as the resulting sample sizes.

Sample Design

The current design is a probability sample stratified by form type and either by 1) size of total assets alone or 2) size of total assets and a measure of income. Form 1120 returns are

Figure E. Total Number of Corporation Tax Returns: Population and Sample Counts, Tax Years 2019–2022

Form type Tax year
Form type 2019 2019 2020 2020 2021 2021 2022 2022
Form type Population Sample Population Sample Population Sample Population Sample
Form type (1) (2) (3) (4) (5) (6) (7) (8)
1120
1120-S
1120-L
1120-PC
1120-RIC
1120-REIT
1120-F
1,729,901
5,153,355
485
16,231
16,582
3,991
51,998
60,713
40,333
485
3,630
9,990
3,064
6,675
1,743,557
5,194,325
450
17,206
18,710
4,414
53,201
61,969
42,501
450
3,888
11,966
3,517
6,914
1,817,159
5,506,634
479
17,568
18,641
4,801
56,696
70,103
48,426
479
4,174
12,310
3,739
7,501
1,800,564
5,638,658
425
18,537
17,522
6,218
58,392
67,930
43,474
425
3,424
11,138
4,831
7,010
Total 6,972,543 124,890 7,031,863 131,205 7,421,978 146, 732 7,540,316 138,232

Bertrand Überall and Nicholas Mountjoy were responsible for the sample design and estimation of the SOI 2022 Corporation Statistics Program under the direction of

Tamara Rib, Chief, SOI Program Support, Statistical Services Branch.

7

2022 Corporate Tax Returns Complete Report Description of the Sample and Limitations of the Data

stratified by size of total assets and size of “proceeds,” which is the measure of income for this form. Size of proceeds is defined as the larger of the absolute value of net income (or deficit) or the absolute value of “cash flow,” which is the sum of net income, several depreciation amounts, and depletion. Form 1120-S is stratified by size of total assets and size of ordinary income. SOI stratified all other 1120-series forms (1120-L, 1120-PC, 1120-RIC, 1120-REIT, and 1120-F) by size of total assets only.

SOI began the design process with projected population totals derived from IRS administrative workload estimates, adjusted using the distribution by population strata from previous survey years. Using projected population totals by sample strata, SOI carried out an optimal allocation based on strata standard errors to assign sample sizes to each stratum such that the overall targeted sample size was 140,000 returns for TY 2022, significantly larger than the TY 2021 target. Mathematical statisticians selected a Bernoulli sample independently from each stratum, with sampling rates ranging from 0.25 to 100%. The total realized sample for 2022, including inactive and noneligible corporations, is 138,232 returns.

Sample Selection

The IRS Kansas City and Ogden Submission Processing Centers process all corporate returns to determine tax liability before transmitting the data to the BMF. After any error correction, these returns are said to “post” to the BMF, which serves as the SOI sampling frame. SOI selects the sample on a weekly basis.

Sample selection for TY 2022 occurred over a 24-month period, from July 2022 through June 2024. SOI requires a 24month sampling period for two reasons. First, nearly 5.6% of all corporations use non-calendar year accounting periods. To capture these returns, the TY 2022 statistics include all corporations filing returns with accounting periods ending between July 2022 and June 2023. Second, many corporations, including some of the largest corporations, request filing extensions, which generally extend the filing deadline by 6 months. This combination of non-calendar year accounting periods and filing extensions means that the last TY 2022 returns the IRS received had accounting periods ending in June 2023, and had to be filed by September 2023. However, taking into account the filing extensions, these returns could have been filed as late as April 2024 and still be considered timely. To account for the normal processing time, the sample selection process remained open for the TY 2022 study until the end of June 2024. In addition, SOI adjusted its processes because some significant returns became available for SOI processing later because of COVID-19 related processing adjustments in the IRS Submission Processing Centers.

Each tax return in the survey population is assigned to a stratum and becomes subject to sampling. Each filing

corporation has a unique Employer Identification Number (EIN). An integer function of the EIN, called the Transformed Taxpayer Identification Number (TTIN), is computed. The number formed by the last four digits of the TTIN is a pseudorandom number. A return for which this pseudo-random number is less than the sampling rate multiplied by 10,000 is selected for the sample.

The algorithm for generating the TTIN does not change from year to year. Therefore, corporations selected for the sample in any given year may be selected for the following year, so long as the corporation files a return using the same EIN and is placed into a stratum with the same or higher sampling rate. If the corporation is placed into a stratum with a lower rate, the probability of selection will be the ratio of the second-year sampling rate to the first-year sampling rate. If the corporation files with a new EIN, then the probability of selection will be independent from the prior-year selection [2].

Data Capture

Data processing for SOI begins with information that was already extracted for IRS administrative purposes. More than 100 data items available from the BMF system are checked and corrected (as necessary), and SOI also extracts some 2,500 additional items from the corporate tax returns during processing. This data capture process can take as little as 15 minutes for a small, single-entity corporation filing Form 1120, or up to several weeks for a large, consolidated corporation filing several hundred attachments and schedules with the return. The process is further complicated by several factors:

●Over 2,500 separate data items may be extracted from any given tax return. This often requires constructing totals from various other items elsewhere on the return.

●Each 1120-series form type has a different layout with different types of schedules and attachments, making data extraction less than uniform for the various forms.

●There is not any legal requirement for a corporation to meet its tax return filing requirements by filling in, line by line, the entire U.S. tax return form. Therefore, many corporate taxpayers report financial details using schedules of their own design or using commercial tax preparation software packages.

●A single accepted method of corporate tax accounting does not exist in the United States, but there are several accepted “guidelines,” which can vary by geographic location. SOI staff attempt to standardize these differences during data abstraction and editing.

●Different companies may report the same data item, such as other current liabilities, on different lines of the tax form. SOI staff also attempt to standardize these differences.

8

Description of the Sample and Limitations of the Data 2022 Corporate Tax Returns Complete Report

Each tax year, to help staff overcome these complexities and differences in taxpayer reporting, SOI prepares detailed instructions for the editing units at the IRS Submission Processing Centers. For TY 2022, these instructions covered standard and straightforward procedures and instructions for addressing data exceptions.

Data Cleaning

SOI staff enter data from the corporate tax returns selected for the sample directly into the database. In this context, the term “editing” refers to the combined interactive processes of data extraction, consistency testing, and error resolution. SOI runs hundreds of tests to check for inconsistencies, and they include identifying:

●Impossible conditions, such as incorrect tax data for a particular form type.

●Internal inconsistencies, such as items not adding to totals.

●Questionable values, such as a bank with an unusually large amount reported for cost of goods sold and/or operations.

●Improper sample class codes, such as when a return has $100 million in total assets but was selected as though it had $1 million because the last two digits of the total assets were keyed in as cents.

Data Completion

In addition to the tests previously mentioned, SOI addresses missing data items and identifies returns to be excluded from the tabulations. The data completion process focuses on these issues.

Beginning with the TY 2012 sample, the criteria for imputing balance sheets for returns with incomplete balance sheets changed significantly. Now, only the largest returns with incomplete balance sheets are subject to SOI’s balance sheet imputation procedure. As a result, the number of returns with imputed balance sheets will be negligible, and SOI will perform imputation on an ad hoc basis only.

SOI uses various methods to impute data for some certainty returns that were unavailable for editing, depending on the information available at the time the return needed to be completed for the sample. These corporations are identified from the previous year’s sample using a combination of assets and receipts. Additional corporations may need to be identified to ensure industry coverage. SOI uses electronically filed data for those corporate returns selected for the sample that were unavailable for statistical processing. For TY 2022, there were 40 returns that met these criteria. For some returns that were not selected for the sample, if the current tax return was not located and other current tax data were not available, then SOI used data from the previous year’s return, with any necessary adjustments for tax law changes.

The data completion process also includes identifying returns not eligible for the sample because the BMF may have duplicate and other out-of-scope returns. These returns include those filed by nonprofit corporations, returns having neither current income nor deductions, and prior-year tax returns. Additionally, amended or tentative returns, nonresident foreign corporations having no effectively connected income with a trade or business located in the United States, fraudulent returns, and returns filed by tax-exempt corporations are not eligible for the sample. Figure F displays the number of inactive sampled returns excluded from the tabulations, as well as the percentages of the total sample size they represent for TY 2019 through TY 2022.

Type of inactive return

No income or

poration Tax Returns: Number of pled Returns for Tax Years 2019–2022
Tax year Tax year Tax year Tax year
2019 2020 2021 2022
(1) (2) (3) (4)
2,602
6,960
2,733
8,235
2,536
10,630
2,837
9,886
9,562 10,968 13,166 12,723
7.69 8.41 9.19 9.22

*Includes duplicate returns (returns that appear more than once in the sample) and prior-year returns.

Figure G provides estimates of the number of active corporations by form type for TY 2019 through TY 2022. For Forms 1120-L and 1120-PC, these estimates may differ from the population counts in Figure E due to changes made during the data capture and data cleaning processes.

Form type

poration Tax Returns: Estimated tive Returns for Tax Years 2019–2022
Tax year Tax year Tax year Tax year
2019 2020 2021 2022
(1) (2) (3) (4)
1,477,196
4,940,351
525
15,589
15,164
3,885
21,037
1,451,658
4,892,722
475
15,870
15,705
4,160
21,540
1,509,261
5,120,552
461
16,155
17,013
4,597
22,692
1,514,763
5,266,702
455
17,195
17,245
5,674
23,687
6,473,747 6,402,130 6,690,732 6,845,719

NOTE: Detail may not add to total due to rounding. *Foreign Insurance Companies file on Forms 1120-L and 1120-PC, but are counted in Form 1120-F, Table 10.

Estimation

SOI bases the estimates of the total number of corporations and associated variables produced in this report on weighted sample data using either a one-step or two-step process, depending on the filed form type. Under the onestep process, SOI assigns a weight for the return, which is the reciprocal of the realized sampling rate, adjusted for unavailable returns, outliers, weight trimming, and any other

9

2022 Corporate Tax Returns Complete Report Description of the Sample and Limitations of the Data

necessary adjustments. SOI used these weights, referred to as the “national weights,” to produce the estimates published in this report for Forms 1120-F, 1120-L, 1120-PC, 1120-RIC, and 1120-REIT, as well as Forms 1120 and 1120-S returns that were sampled with certainty.

The two-step process is used to improve the estimates by industry for returns filed using either Form 1120 or Form 1120-S that were not selected in self-representing strata. The first stage of the two-step process is to assign an initial weight for the return as previously described. The second stage involves post-stratification by industry and sample selection class. SOI uses a bounded raking ratio estimation approach to determine the final weights because certain post-stratification cells may have small sample sizes [3]. SOI used these final weights for these forms to produce the aggregated frequency and money amount estimates that are published in this report.

Data Limitations and Measures of Variability

SOI uses several extensive quality review processes to improve data quality. This starts at the sample selection stage with weekly monitoring to ensure the proper number of returns is selected, especially for the certainty strata. These processes continue using consistency testing through the data collection, data cleaning, and data completion procedures. Part of the review process includes extensive comparisons between the sample year (TY 2022) and prior-year (TY 2021) data. SOI designed each processing stage to ensure data integrity.

Sampling Error:

Since the TY 2022 estimates are based on a sample, they may differ from population aggregates, which were compiled from a complete census of all corporate income tax returns. The TY 2022 sample is one of many possible samples that could have been selected under the same sample design. Estimates derived from one possible sample could differ from those derived from another sample or from the population aggregates. The deviation of a sample estimate from the average of all possible similarly selected samples is called the sampling error.

The standard error (SE), a measure of the average magnitude of the sampling errors over all possible samples, can be estimated from the realized sample. The estimated standard error is usually expressed as a percentage of the value being estimated. This is called the estimated coefficient of variation (CV) of the estimate, and it can be used to assess the reliability of an estimate. The smaller the CV, the more reliable the estimate is deemed to be.

SOI calculates the estimated coefficient of variation of an estimate by dividing the estimated standard error by the estimate itself, and then taking the absolute value of this ratio. Table 1 (see related Complete Report tables) shows the estimated coefficients of variation by industrial groupings for the estimated number of returns as well as selected money amounts.

10

The estimated CV, CV(X), can be used to construct confidence intervals for the estimate X. The estimated standard error, which is required for the confidence interval, must first be calculated. For example, the estimated number of companies in the manufacturing sector with net income and the corresponding estimated CV can be found in Table 1 and used to calculate the estimated standard error:

SE(X) = X • CV(X)

= 138,867 x 3.92/100

= 5,444

A 95% confidence interval for the estimated number of returns in manufacturing is constructed as follows:

X ± 2 • SE(X)= 138,867 ± (2 x 5,444)

= 138,867 ± 10,888

The interval estimate is 127,979 returns to 149,755 returns. This means that if all possible samples were selected under the same general conditions and sample design, and if an estimate and its estimated standard error were calculated from each sample, then approximately 95% of the intervals from two standard errors below the estimate to two standard errors above the estimate would include the average estimate derived from all possible samples. Thus, for a particular sample, it can be said with 95% confidence that the average of all possible samples is included in the constructed interval. This average of the estimates derived from all possible samples would be equal to or near the value obtained from a census.

Nonsampling Error:

In addition to the sampling error, a nonsampling error can also affect the estimates. Nonsampling errors can be classified into two groups: random errors—whose effects may cancel out, and systematic errors—whose effects tend to remain somewhat fixed and result in bias.

Nonsampling errors include coverage errors, nonresponse errors, processing errors, or response errors. The inability to obtain information for all sampled returns, differing interpretations of tax concepts or taxpayer instructions, inability to provide accurate information at the time of filing (data are collected before auditing), and inability to obtain all tax schedules and attachments may cause these errors. These errors may also be caused by data recording or coding errors, data collecting or cleaning errors, estimation errors, and failure to represent all population units.

Description of the Sample and Limitations of the Data 2022 Corporate Tax Returns Complete Report

Coverage Error:

Coverage errors in the SOI corporation data can result from the difference between the time frame for sampling and the actual time needed for filing and processing the returns. Since many of the largest corporations receive filing-period extensions, they may file their returns after the closing date for sample selection, as was explained before in the Sample Selection description. However, any of the largest returns found are added into the file until the final file is produced.

Coverage problems within industrial groupings in the SOI Corporation study may result from the way some consolidated returns are filed. The Internal Revenue Code (IRC) permits a parent corporation to file a single return, which includes the combined financial data of the parent and its subsidiaries. These data are not separated into the different industries but are entered into the industry with the largest receipts. Thus, there is undercoverage of financial data within certain industries and overcoverage in others. Coverage problems within industries present a limitation on any analysis of the sample results.

Nonresponse Error:

There are two types of nonresponse errors: unit and item. Unit nonresponse occurs when a sampled return is unavailable for SOI processing. For example, other areas of the IRS may have the return at the time it is needed for statistical processing. These returns are termed “unavailable returns.”

Item nonresponse occurs when certain items are unavailable for a return that was selected for SOI processing, even if the return itself is available. An example of item nonresponse would be an item missing from the balance sheet, even though other items have been reported.

Processing Error:

Errors in recording, coding, or processing the data can cause a return to be sampled in the wrong sampling class. This type of error is called a misstratification error. An example of how a return might be misstratified: a corporation files a return with total assets of $100,000,023 and net income of $5,000. A processing error causes the last two digits of the total assets to be keyed in as cents, so that the return is classified according to total assets of $1,000,000.23 and net income of $5,000.00. The return would be misstratified according to the incorrect value of the total assets stratifier. To adjust for misstratification errors, only returns selected in a noncertainty stratum that actually belonged in a certainty stratum were moved to this certainty stratum.

Response Error:

Response errors are due to data being captured before audit. Some purely arithmetical errors made by the taxpayer are corrected during the data capture and cleaning processes. Because of time constraints, SOI does not incorporate adjustments to a return during audit into the file.

Corporation Income Tax Returns for Statistics of Income, 1951 to Present,” 1984 Proceedings of the Section on Survey Research Methods, American Statistical Association, pp. 437–442.

[2] Harte, J. M. (1986), “Some Mathematical and Statisti cal Aspects of the Transformed Taxpayer Identification Number: A Sample Selection Tool Used at IRS,” 1986 Proceedings of the Section on Survey Research Methods, American Statistical Association, pp. 603–608.

[3] Oh, H. L., and Scheuren, F. J. (1987), “Modified Raking

Ratio Estimation,” Survey Methodology, Statistics Canada, Vol. 13, No. 2, pp. 209–219.

References

[1] Jones, H. W., and McMahon, P. B. (1984), “Sampling

11

Get a plain-English answer with a citation back to this text.

Ask AI about this code
▸Contents — 0925 Publ 16 (PDF)

GoCodebook provides public access, search, citation, multilingual explanation, and practical interpretation of legally adopted building regulations. It is not a substitute for the official ICC or California code publications.