The objective of the literature review presented in this chapter is to identify and assess relevant methods researchers apply to address misreporting and incompleteness in crash data, and to evaluate the methodsʼ strengths, weaknesses, and appropriateness in reducing the misreporting of motor vehicle crashes involving impaired and distracted driving.
Misreporting in motor vehicle crash data can be an issue of reporting (overreporting or underreporting) or of recording (omitted, incomplete, or inaccurate records, which can lead to the misclassification of crashes). A detailed analysis into the causes of errors in crash data reporting can be found in Ahmed, Sadullah, and Yahya (2019), wherein the authors identified 26 causes of errors, 12 of which were related to reporting and 14 of which were related to recording. Crash underreporting is the most common misreporting issue and therefore has been studied by several researchers (Alsop and Langley 2001; Salifu and Ackaah 2012; Watson, Watson, and Vallmuur 2015).
The causes of misreporting in crash data are manifold and can be beyond the control of a reporting agency. For example, police may not be notified of a crash because the parties involved opt for a private settlement for insurance purposes, the crash involves only one vehicle, or the statutory damage minimums for reporting are not met. Additionally, crashes may not be reported if no obvious injury is observed immediately after the crash (Amoros, Martin, and Laumon 2006). Generally, the filing of crash reports correlates with injury severity level and road user type. Previous studies have found that fatal and serious injury crashes are more likely to be reported to the police than minor injury and property-damage–only crashes (Elvik and Mysen 1999; Alsop and Langley 2001; Amoros, Martin, and Laumon 2006; Abay 2015). Among road users, the reporting rate is lowest when crashes involve motorcyclist and cyclist injuries (Salifu and Ackaah 2012; Watson, Watson, and Vallmuur 2015). Among different crashes, those involving impaired and distracted driving have higher rates of misreporting.
Impaired driving can involve alcohol, drugs, or both. The underreporting of alcohol and drug involvement in crashes likely stems from a few factors. Universal drug testing of all drivers is expensive to implement, and testing seriously injured drivers is challenging. Furthermore, emergency departments often do not test for or report alcohol or drug impairment of drivers injured in crashes because many emergency departments have limited time and resources. Additionally, some medical staff may be concerned that citing alcohol involvement would allow health insurers to deny coverage (Miller et al. 2012). Patient privacy laws create a further barrier to information sharing with law enforcement when testing is completed in the hospital setting.
Similar to impaired driving, distracted driving can also be underreported. According to NHTSA (2020), distracted driving is likely to be underreported in crash data because of three
potential causes. First, drivers are not likely to self-report distracted driving (i.e., texting) because they fear receiving a citation. Second, if a distracted driver loses their life in a crash, officers may not be able to report distracted driving as a cause of the crash. Even when witnesses are available, their accounts may not be accurate or detailed enough to support a judgment by the officer. Third, police crash reports may be out of date and unable to capture data from the latest technologies, and consequently lack the detailed fields and codes needed to document distraction.
Crash records linked with hospital records and using methods such as the Crash Outcome Data Evaluation System (CODES) serve as the primary data source to examine the misreporting of alcohol-impaired crashes. CODES is a state-based program designed in 1992 by NHTSA. The program used probabilistic linkage to combine data from vehicle crash reports with emergency department and inpatient hospital records. In 2013, NHTSA transitioned CODES to full state-level responsibility (NHTSA 2015), and several states have maintained the effort or started similar data-integration projects on their own.
NCHRP Web-Only Document 302: Development of a Comprehensive Approach for Serious Traffic Crash Injury Measurement and Reporting Systems surveyed states and reported the identifiers for linkage with EMS data, emergency department data, hospital discharge data, trauma registry data, vital records data, and roadway inventory data (Flannagan, Rupp, and Mann 2015). It was also noted that for linking crashes to medical data, available identifiers and matching variables differ among states (NHTSA 2015). Miller et al. (2012) used capture–recapture methods with CODES data to estimate that 7.5 percent of drivers in non-fatal crashes and 12.9 percent of non-fatal crashes were alcohol involved. However, Miller et al. found only 44 percent of drivers noted as alcohol involved in the hospital record were coded as such by law enforcement on the crash report. Reporting by LEOs aligned better with hospital records as the severity of crash victimsʼ injuries increased.
Other studies have examined only the extent of underreporting alcohol-involved crashes. Orsay et al. (1994) examined drivers involved in crashes from two trauma centers, finding 31.2 percent of drivers had BAC levels at or above 0.10 g/dL and a further 15 percent of drivers had positive drug screens for an overall impairment rate of 46.2 percent. However, only 16.5 percent of drivers were cited for driving under the influence. In a similar analysis from Maryland of drivers with positive BAC levels recorded by the hospital, LEOs did not record alcohol involvement for 42 percent of drivers, including 39 percent who had BAC levels above 0.08 g/dL (Miller et al. 2012).
Subramanian (2002) reported on the transition to multiple imputation as a new method for estimating missing BAC in the Fatality Analysis Reporting System (FARS). As described in the report, “the new methodology improves on the current model by imputing specific values of BAC across the full range of possible values rather than estimating probabilities.” The methodology replaces missing BAC with 10 simulated values, which can then be used to calculate valid statistics. The report found that multiple imputation estimates were 2 percent higher than the estimates calculated using the previous method; notwithstanding this increase, the overall trend of alcohol involvement was similar for both methods.
Definitions, drugs considered, policies, practices, and laws regarding drug-impaired driving and toxicology testing vary widely among states. In a 2016 report, Advancing Drugged Driving Data
at the State Level, Arnold and Scopatz noted several issues leading to underreporting drug-impaired driving. Underreporting issues included the cost of conducting a drug panel, insufficient officer training to detect drug impairment, a lack of sensitive field tests for detecting drug impairment, the high cost of roadside toxicology tests, and the challenges of implementing roadside toxicology tests. Further, researchers note that most states do not distinguish among driving under the influence of alcohol, drugs, or both. Similarly, most states do not apply enhanced penalties for drug impairment in addition to alcohol impairment, which creates a disincentive to investigate drug use when a crash-involved driver has a BAC level at or above 0.08 g/dL. Lastly, crash reports have limited capacity to capture toxicology results.
In its report Undercounted Is Underinvested: How Incomplete Crash Reports Impact Efforts to Save Lives, the NSC (2017) found that officers have difficulty detecting prescription drugs or OTC drugs and that alcohol impairment may be underreported because of the difficulty in observing impairment from alcohol concentrations less than 0.08 g/dL. Arnold and Scopatz (2016) noted that drugged driving may be overreported when drugs may be detectable in a crash victimʼs system after impairment ends. In short, polydrug use and a mix of alcohol with other drugs can complicate the picture for law enforcement, testing, and analysis.
NHTSAʼs report Alcohol and Drug Prevalence Among Seriously or Fatally Injured Road Users (2022) was the largest research effort in the United States that conducted independent toxicological analysis of road users. NHTSA selected seven Level 1 trauma centers to participate. Duration of data collection varied by trauma center, ranging between 9 and 23 months. In total, 7,279 roadway users met the studyʼs inclusion criteria. Centers reported that 55.8 percent of injured or killed roadway users (drivers, pedestrians, bicyclists, and passengers) tested positive for one or more drugs. Cannabinoids were at 25.1 percent positive, alcohol at 23.1 percent, stimulants at 10.8 percent, and opioids at 9.3 percent. The report also found that 19.9 percent tested positive for two or more categories of drugs.
Distracted driving is largely believed to be underreported in crashes (Stutts et al. 2001). One reason for unavoidable underreporting of distracted driving is the difficulty LEOs have in detecting or observing distraction (Ranney 2008; National Traffic Law Center 2017). Officers are unlikely to report distraction unless there is direct evidence (Ranney 2008). Evidence can also be unavailable because of contamination from life-saving efforts of emergency medical, fire, or rescue personnel (National Traffic Law Center 2017). Further, officers may be less likely to report observed distracted behaviors that are legal, such as talking on a cell phone, versus actions that are explicitly illegal (NSC 2017). Barring direct evidence, officers are reliant on driver or occupant statements. Ranney (2008) notes that “drivers are understandably reluctant to admit that they were engaged in a secondary task, particularly if that involvement may have contributed to the crash.” Further, and specifically regarding distracting use of electronic devices, it is costly and time consuming to investigate device use at the precise time of a crash (e.g., subpoena cell phone records and align calls and texts with the moments leading up to a crash).
Sources of distraction that do not lend themselves to some level of documentation can be impossible to investigate after the fact. In its report Distracted Driving Prevalence Data Sources, Challenges and Technological Solutions, the National Distracted Driving Coalition (NDDC) describes data collection barriers to reporting distracted driving. Barriers include police being unable to identify distracted driving because the drivers do not admit to it or are killed in the crash, limitations of existing crash report forms, and data systems not being structured to capture and query distraction-related data. Self-reported data are limited by respondents not admitting to driving distracted because it is not a socially desirable response. Additionally, drivers may not
recall the timing of events when questioned (NDDC 2022). While hospital data may note that the patient was involved in a vehicle crash as the reason for admission, information regarding crash factors is recorded in narrative notes and as a result is difficult to query (NDDC 2022).
In its report Impaired and Distracted Driving: Data Comparison, the Traffic Injury Research Foundation (TIRF) estimated that 20 percent to 30 percent of fatal collisions in North America can be attributed to distracted driving, as supported by self-reported data from 2020. Self-reports showed that 13.3 percent of Canadians talk on the phone while driving and 11.2 percent text while driving (TIRF 2021). The TIRF also covered limitations to reporting distracted driving. Those limitations included some crash report forms not including distraction as a driver condition, officers being limited on the number of driver conditions they are able to enter on forms, officers being unable to determine distraction in fatal crashes, and distraction behavior being included in narrative sections of reports and thus hard to query (TIRF 2021).
Crash report forms and crash databases are primary data sources for road safety research. Imprialou and Quddus (2019) reviewed the current literature on the state of crash data quality and found the most serious data quality issues to be location and time inaccuracies, difficulty linking crash and traffic data, misclassification of crash severity, inaccurate or incomplete information on demographics of involved persons, and crash-contributing factors. Crash reports include a fixed list of contributing factors from which officers can select. Research findings showed that officers are likely to report the minimum permitted number of contributing factors for several reasons, including officer level of training, time restrictions for completing the report, and the complexity of reporting crash mechanisms using a generic, pre-designed form. Additionally, findings describe how contributing factors may suffer from lack of objective and generally accepted descriptors, as the perceptions of drivers and police officers may differ.
In the United Kingdom, Rolison (2020) investigated crash report forms and the classification of contributory factors. This report found that police officersʼ categorical perceptions could lead to misunderstanding and misreporting. Rolison (2020) used hierarchical clustering to identify an optimal category structure based on police officer perceptions and found that perception of contributing factors was influenced by imposed categories. Dissimilar factors could be perceived as similar when classified in the same category.
In a separate study, Rolison et al. (2018) investigated the cause of crashes by comparing the views of police officers with the views of the general driving public. They found that driver age and gender influenced the ratings of both police and the public about the most likely causes of crashes. The public indicated a higher likelihood of distraction across scenarios. Distraction was rated as less likely to contribute to male driver crashes than to female driver crashes. Drugs and alcohol were rated as more likely to contribute to crashes by male drivers than female drivers and to crashes with young drivers. Police officers perceived drugs or alcohol as a likely contributor to crashes by middle-aged and young drivers.
Researchers have used various methods to study and address the issue of misreporting in crash data. Although most studies were unsuccessful in identifying effective methods to quantify the misreporting of certain crash types, the methods used may be applicable for studying the misreporting of impaired and distracted driving crashes. Identified methods, including linking crash records with alternative data sources, capture–recapture, text mining of crash narratives, crash count underreporting, and injury severity underreporting are compared in Table 1.
The column headers of the table are Method, Purpose, Data, Strengths, and Limitations. The data given in the table row-wise are as follows:
Row 1: Linking Crash Records with Alternative Data Sources; Capture missing information in crash data from alternative data sources and link with crash records to obtain better understanding of crash pattern, count, injury severity, under-reporting rate, misclassification, and so forth; Crash data and alternative data sources, such as hospital records, road trauma registry data, emergency medical service data, and survey-based data; Estimate actual crash count, Estimate the extent of under-reporting based on different categories, such as injury severity level, road user type, and crash type. Calculate under-reporting rate, Unbiased analysis results can be obtained by using the linked data as missing data in the crash record, Can capture additional information not recorded in crash data. Data linkage is possible only if properly maintained alternative data sources are present. Alternative data sources with unique identifiers are needed for proper linkage. Coding or reporting errors in alternative data sources can hamper linkage accuracy. The accuracy of the linkage depends on the linking approach: probabilistic or deterministic.
Row 2: Capture Recapture: Using a sampling method to estimate the true population of crash data; Crash data and alternative data sources; Can estimate unknown population size through sampling or true crash count, The under-reporting rate can be estimated after estimating the true population; Sampling accuracy depends on alternative data sources; Probabilistic approach-based mismatch in alternative data sources reduces accuracy, Further statistical analysis is needed to estimate the factors contributing to misreporting.
Row 3: Text Mining of Crash Narratives: Identify crash types based on analysis of crash narratives; Crash narrative in a crash report; Can classify crashes based on valuable information extracted from crash narratives; Can recover misclassified crashes; Automate analysis of crash narratives with increased efficiency, Model performance can be poor if the text data are noisy; An arbitrary threshold value is needed for classification; Careful interpretation of results is necessary as crash narratives of the same crash type may vary based on recording authority; Need alternative sources of data to validate the results.
Row 4: Crash Count Under-reporting Model: Estimate the under-reporting rate, identify contributing factors to under-reporting, and tackle the impact of under-reporting on crash frequency, Crash data, Crash data are enough for analysis; Estimate the bias of factors contributing to crash under-reporting. Need alternative sources of data to validate the results.
Row 5: Injury Severity Under-reporting Model, Estimate injury severity under-reporting and reporting bias of injury severity, tackle the impact of under-reporting in injury severity, Crash data, Analysis is based only on crash data; Estimates the under-reporting of injury severity in police records; Estimates the factors contributing to injury severity under-reporting; Adjusts the bias in parameter estimates associated with misclassification, Cannot estimate the true crash counts; Need alternative sources to validate the results.
Crash data can be linked with alternative data sources as a method to address misreporting issues with impaired and distracted driving. Many previous studies have used alternative data sources such as hospital records, trauma registry data, emergency department data, and survey data. The studies are categorized in this section by data source.
Several studies have linked hospital records with crash records to measure the underreporting rate of impaired driving (Alsop and Langley 2001; Yannis et al. 2014; Kamaluddin, Abd Rahman, and Várhelyi 2019). Alsop and Langley (2001) linked hospital records and police records using AutoMatch, a commercially available record linkage software that uses probabilistic methods. Before joining the data with crash records, the authors selected hospital records based on certain codes in the hospital data that represented motor vehicle traffic crashes, primary diagnosis of injury, vehicle occupant, and first admission (excluding readmissions). Of all selected records available in the hospital file, 63 percent were linked to crash records. The reporting rate was defined as the number of linked hospital records divided by the total number of hospital records. Chi-square tests were used to assess the statistical differences in the reporting rates between different groups, and a multivariate stepwise logistic regression was used to examine the effects of different factors on underreporting. The research found that reporting rates were higher for drivers when compared with passengers, and there were statistically significant differences in the reporting rates of types for road users, crash types, and number of vehicles involved. The research also found greater likelihood of reporting as the injury severity increased.
A Malaysian research study used unique personal identifiers, such as individual identification numbers, given names, and surnames, to link crash data in police and hospital databases (Kamaluddin, Abd Rahman, and Várhelyi 2019). Also, some secondary attributes (specifically date of crash, age, and gender) were used for verification purposes. Data linkage was performed using the MSSQL relational database management system software, and both deterministic and probabilistic approaches were used to obtain maximum linkages. The deterministic approach was based on exact matches of unique personal identifiers. The probabilistic approach used a combination of asterisks and wildcards to substitute characters in the personal identification number, given name, or surname to ensure that alternative spellings and numbers were included. The study findings concluded that 7,625 persons involved in crashes were registered in the hospital database, while only 362 were recorded by police. The study also found a large disparity between police and hospital records: The hospital database contained 86 percent of police-reported cases; however, the police database contained only 4.1 percent of hospital cases. Moreover, the matching rate between police and hospital records was found to be proportional to the level of injury severity.
Yannis et al. (2014) used underreporting coefficients to model the extent and variation of injury underreporting in eight European countries. The study adopted a common approach for each country, following a framework of linking the crash database (usually maintained by the police) with a medical database (from hospitals) to identify all common records and copy details from the medical record of each linked crash to the corresponding record in the police database. The linkage method is probabilistic, so uncertainty ranges are involved. Also, for ambiguous records (e.g., French and Greek studies), manual linkage was implemented to minimize the uncertainty ranges. After the data were linked, log-rate models were developed to estimate the combined effects of country, road user type, police severity score (serious or slight injury), and MAIS score (Maximum Abbreviated Injury Scale score) on underreporting. The authors recommended that, in addition to improving police records, improved recording of injuries at hospitals and other medical facilities will address the problem of underreporting and misclassification of injuries.
Guo et al. (2007) investigated the effects of misclassifying seat belt and alcohol use on the odds ratio of injury. They also studied the changes to estimated medical costs after data linkage of hospital and crash records. First, the authors fitted a logistic regression model using Nebraska CODES data from 1996 to 1997 to evaluate the effect of seat belt and alcohol use on injury severity. Next, they fitted the logistic regression model again after using a self-developed SAS computer program to adjust proportions for the misclassification of seat belt and alcohol use. The misclassification rates were estimated from a data set of hospital records merged with police records. The estimated misclassification rate of 20 percent might not be precise because of the small size of the merged Nebraska police and hospital data set.
In Amoros, Martin, and Laumon (2006) and Amoros et al. (2007), police-reported crash data were linked with road trauma registry data to estimate the degree of injury severity misreporting. The studies found that for New Injury Severity Score (NISS) of 1 to 3, police reported 2.8 percent as seriously injured. For a NISS of 2 to 8, police reported 17.1 percent as seriously injured. For a NISS of 9 to 15, police reported 46.0 percent as seriously injured. For a NISS of 16 to 24, police reported 65.4 percent as seriously injured. Lastly, for a NISS of 25 to 75, police reported 80.3 percent as seriously injured.
In a similar study, a multivariate analysis of the probability of police severity misclassification was performed. Tsui et al. (2009) determined the extent of injury severity misclassification by measuring the degree of discordance between police reports and hospital records on injury severity after linking crash records with trauma records from the regional hospital. For Injury Severity Scores of 1 to 15, police classified 80.2 percent as slight injury and 19.8 percent as serious injury. For Injury Severity Scores of 16 to 75, police classified 5.3 percent as slight injury and 94.7 percent as serious injury. In a third study, researchers linked crash data with emergency room data to find the reporting propensity of drivers using a binary response model (Abay 2015). They concluded that unaccounted reporting bias in crash data can mislead road safety policies.
Miller et al. (2012) linked data from police crash reports and hospital inpatient and emergency department discharges to identify alcohol-involved injury crashes, evaluate the accuracy of police and hospital reports for alcohol-involved crashes, and learn how reporting varies among different states. They found large divergence in reporting, with police reporting alcohol involvement for 44 percent of persons the hospital reported as alcohol involved and hospitals reporting alcohol involvement for 33 percent of the persons identified as alcohol involved by police. They concluded that police and hospitals need to improve communication about alcohol involvement.
Castle et al. (2014) investigated the underreporting of alcohol involvement on death certificates by linking national Multiple Cause of Death data with FARS data. The ratio of reporting alcohol involvement on death certificates was computed as the prevalence of any mention of alcohol-related conditions among motor vehicle traffic deaths in Multiple Cause of Death data, divided by the prevalence of deceased drivers with BAC test results of 0.08 percent or greater in FARS. Researchers concluded that the comparison between FARS and Multiple Cause of Death data showed large discrepancies in reporting alcohol involvement and considerable variation among states for underreporting.
Multiple alternative data sources were linked in a study by Watson, Watson, and Vallmuur (2015). Crash data were provided from the Queensland Road Crash Database. Alternative data sources linked with crash data included the Queensland Hospital Admitted Patient Data Collection, the Emergency Department Information System, and Queensland Injury Surveillance Unit data. Researchers calculated a discordance rate to find the underreporting rate and did a statistical analysis to find out how different characteristics influenced the discordance rate.
Yang et al. (2019) found the factors that influence the underreporting of minor crashes using self-reported survey data. A retrospective survey was conducted in both Beijing and Kunming,
China, in June 2013 and July 2013, investigating the crash history of more than 3,000 respondents as well as police-reported crash data. Underreported rates for automobile-to-automobile crashes were found to be 56 percent; automobile-to-non-motorized vehicle crashes, 77 percent; and non-motorized-to-non-motorized vehicle crashes, 94 percent. Because respondents might not have remembered some details of past crashes, such as time, weather, road geometry, road surface condition, and lighting condition, it is likely that not all information was collected. Also, surveys can be time consuming and expensive. Another limitation of this data collection method is that crash severity is reported by only one party involved; therefore, the injury severity of the other party may not be accurate, especially in crashes reported by automobile drivers.
Similarly, Salifu and Ackaah (2012) conducted surveys at hospitals and among drivers to generate relevant alternative data, which were then matched against records in police crash data files and the official database to estimate the overall shortfall (i.e., underreporting) in the official crash statistics in Ghana over 8 years. They found that “shortfalls” came mainly from non-reporting and underreporting. The level of non-reporting varied by crash severity, with 57 percent for property-damage–only crashes, 8 percent for serious injury crashes, and 0 percent for fatal crashes.
Vissers, Houwing, and Wegman (2018) used survey data to identify critical elements that cause alcohol-related crashes to be underreported. The study identified existing methods that could be used to counter underreporting and improve the reliability of the official crash statistics on drunk driving. These included methods such as police testing 100 percent of all road users involved in a crash for alcohol, harmonizing definitions of alcohol-related road casualties, and establishing legal limits on alcohol for pedestrians and bicyclists. The online survey in the study was carried out in 45 countries and included questions regarding recording methods, data sources, definitions, and current legislation related to alcohol-related crashes.
Medury et al. (2019) studied non-motorized traffic safety concerns in and around three university campuses by comparing police-reported crash data with traffic safety information crowdsourced from the campus communities. They found that police-reported crashes underrepresent non-motorized safety concerns, while the self-reported crash results reported a wide variety of collisions not involving automobiles. Their findings revealed that self-reported crashes did not have significant overlap with police-reported crashes. The authors acknowledged that crowdsourcing crash data can be biased because of recall bias (i.e., the inability to recollect older details), participation bias (certain stakeholders having a greater interest in the survey or better access to the internet), and reporting bias (the inability to verify the claims made in the surveys). Self-reported crash data may not be a viable alternative to police-reported data, as the information provided by those involved in the crash cannot be verified by the other parties involved or by trained professionals.
The sampling method capture–recapture is used in several studies to estimate the unknown size of a population (Bos, Reurings, and Derriks 2009; Miller et al. 2012; Janstrup et al. 2016). Janstrup et al. (2016) computed the underreporting rate using the capture–recapture method and a binary logit model estimating the probability that a road crash injury, n, appears in the database, M, given that the same road crash injury appears in the other database. The capture–recapture method estimates based on overlapping records in two independent samples under four assumptions: the population is finite and closed, common records are unambiguously identified, records are independent, and records are homogeneously catchable. The limitation of the study is that although model estimates reveal that reporting to the police and the hospital can be explained by individual characteristics, the data lack information about the reasons for underreporting.
Underreporting and biased reporting issues are explored in Bos, Reurings, and Derriks (2009) based on data sources in addition to crash databases. This study used a capture–recapture method to match crash and hospital data sources and to quantify the underreporting amount and identify reporting bias. They concluded that to quantify underreporting and bias, it is necessary to develop alternative data sources. No investigation into the accuracy of alternative data sources was provided.
Miller et al. (2012) used a capture–recapture statistical model to estimate unreported cases from the extent of reporting overlap. The study estimated the amount of alcohol-involved injury crashes along with police and hospitalsʼ reporting accuracy of those alcohol-involved crashes and examined how reporting varies among states. The study analyzed overlap in alcohol reporting with the limitation that police might not report alcohol involvement for BAC results below the per se limit. They reported that police correctly identified 32 percent of alcohol-involved drivers in non-fatal crashes. Also, hospitals reported 28 percent involvement for emergency department cases and 51 percent involvement for admitted cases. Coding errors, missing data, and non-uniformity of administrative data sets imposed further limitations. A few alcohol-involved cases might have been removed, as police misreported those as drug involved. Furthermore, CODES linkages are probabilistic; the occasional mismatch reduced reporting consistency and capture–recapture accuracy.
By reviewing crash narratives or crash reports by machine learning and text-mining techniques, researchers have tried to improve the misclassification error of misreporting in crash data and classify different crash types (Zhang et al. 2020; Sayed et al. 2021). In the NoisyOR method, the probability of a crash being a specific type is calculated by combining the probability scores of unigrams (single words) and bigrams (two consecutive words) in the narrative. NoisyOR is a probabilistic extension of the logical “or” (Oniśko, Druzdzel, and Wasyluk 2001; Vomlel 2006). If any input has a high probability score (such as a value close to 1), then the combined probability in NoisyOR becomes high. The combined probability in NoisyOR is even higher if more input probabilities are high. To apply a NoisyOR classifier to crash narratives, it is necessary to compute the probabilities of unigrams, bigrams, and trigrams (three consecutive words) and combine these probabilities using the NoisyOR method.
Sayed et al. (2021) identified misclassified work zone crashes using text mining on crash narratives. In this study, a keyword-based text classifier was developed using the NoisyOR combined probability to identify misclassified work zone crashes from the crash narratives of police reports. A manual review of the top 450 cases classified as work zone crashes by the model revealed that 201 were indeed work zone crashes, while the remaining 249 were not. The authors used the NoisyOR method because of its ability to work effectively despite the high level of noise in the unstructured text or crash narratives. This method does not require significant training time, and it is computationally efficient and easier to implement than other text-mining techniques; however, the method does require manual review to verify positive cases.
Zhang et al. (2020) used a text-mining approach for analyzing crash narratives to identify secondary crashes. Because narratives are unstructured and cannot be understood directly by machine learning algorithms, they were first transformed into numerical vectors following a four-step process. Numeric vectors were then fed into four popular machine learning models: logistic regression, random forest, Naïve Bayes, and support vector machine. The logistic regression model outperformed the other three models in terms of overall accuracy and F1 score. A review of the results showed that the model was effective at identifying keywords for secondary crashes.
Novel statistical methods have been proposed by researchers to tackle misreporting, especially underreporting issues. Underreporting in crash count and crash severity modeling is handled by different statistical methods. Statistical models are usually developed under the assumption that data are randomly sampled, and each crash has an equal opportunity of being sampled. However, not all crashes are reported because of varying reporting thresholds and errors in reporting and recording.
Wood, Donnell, and Fariss (2016) formulated a negative binomial underreporting model used to define a reporting probability and multiplied it by the true crash count to find the reported crash count. The underreporting model estimated the probability of reporting to predict underreported crashes. This model was specified as a binary logit model of a set of factors contributing to the underreporting issue. Moreover, when compared with random parameter models that account for unobserved heterogeneity, the negative binomial underreporting model was found to provide better predictions, which proved the effectiveness of the model in crash frequency research. Further, the study validated the underreporting model by comparing predicted underreported crashes with crashes reported to police but not included in the crash database.
A copula regression model was developed to investigate the impact of underreporting on wildlife–vehicle collision (WVC) data analysis (Zou et al. 2019). The study combined a negative binomial WVC model with an underreporting outcome model. For the underreporting outcome model, authors generated a new dichotomous underreporting indicator variable to denote underreporting probability (underreporting: 1; otherwise: 0) and adopted a logistic regression model to analyze the impact of explanatory variables on the underreporting outcome. Then, a bivariate copula approach was used to link the occurrence of WVCs and the underreporting probability. Lastly, using the joint cumulative density function combining occurrence of WVCs and the underreporting probability (i.e., indicator variable), a copula regression model was formulated. From the model output, authors concluded that the proposed copula model was found to be a better alternative to the conventional negative binomial model for modeling underreported WVC data as it could identify hotspots more accurately than the negative binomial–based Empirical Bayes method.
Kumara and Chin (2005) used a Poisson underreporting model to analyze the factors affecting road crash frequency at three-legged signalized intersections. To handle the effect of underreporting, the mean function of Poisson regression was first modified by introducing unobserved heterogeneity (site-specific unobserved factors affecting crash occurrence). Next, a pre-established reporting mechanism was combined with the modified Poisson regression to get the final version of the modified Poisson mean for reported crashes. Lastly, the authors introduced endogenous and exogenous Poisson regression probit models depending on whether there was correlation between variables contributing to crash occurrence and reporting. It was found that the exogenous probit underreporting model is a better representative model to its parent Poisson model and so was used in the study. The model outcome showed that underreporting was affected when proper legal requirements were not imposed for reporting non-injury crashes. Also, the geographic location of the intersections was a factor behind underreporting.
Zeng et al. (2020) developed a Bayesian underreporting conditional autoregressive (CAR) model for analyzing crash frequency to simultaneously deal with both underreporting and spatial correlation issues. In this study, to account for the spatial correlation across adjacent crash sites, a CAR prior term was added into the generalized linear function of the traditional CAR model. Moreover, a residual term normally distributed with zero mean was added to capture the unstructured heterogeneous effects. For estimating underreporting in the CAR model, a crash reporting rate was integrated into the counting process in the form of binary logit (Wood, Donnell,
and Fariss 2016). Thus, authors found the average or mean of the reported number of crashes as the product of true crash count mean and the reporting rate. The results obtained from the developed model were consistent with the findings in existing literature and engineering experience, which implied it was a good alternative for predicting crash frequency.
Sequential binary probit models and ordered-response probit models of injury severity were developed by Yamamoto, Hashiji, and Shankar (2008). In their study, the misclassification problem was treated as outcome-based samples with unknown population shares of the injury severities. Researchers introduced new variables Q(j) as the unknown population share of severity j and applied the pseudo-likelihood function for outcome-based samples to examine the effects of severity underreporting on the parameter estimates. The estimated reporting rate was calculated as the ratio of reported injury outcome to the estimated population share. Research results showed the reporting rate for property-damage–only crashes and possible injury to be 27.5 percent and 2.6 percent, respectively.
Ye and Lord (2011) also treated the misclassification problem as outcome-based samples instead of random samples from the population. In the crash severity model, each crash was weighted by the ratio of the true crash severityʼs population share Q to the sample share H, or the severity share for the observed crashes. The authors used the WESMLE (i.e., the weighted exogenous sample maximum likelihood estimator) to account for specific underreporting conditions with full and partial knowledge of different severity unreported rates. After testing with multinomial logit, ordered probit, and mixed logit models using simulated and observed crash data, the authors concluded that when some information regarding the extent of underreporting issue is known, using a WESMLE produces better results than the maximum likelihood estimator.
Similarly, Patil, Geedipally, and Lord (2012) developed a nested logit model for crash severity analysis with the assumption that outcomes-injury severity levels in the population are not known. In the utility function, the authors introduced a new term, S(i, θ), to represent the sampling probability for the roadway segment I, and this term can contain unknown parameters θ. New parameters could be estimated by a weighted conditional maximum likelihood estimator. The study demonstrated an approach to tackling the underreporting issue within a nested logit model.
Balan and Paleti (2018) deployed a mixed generalized ordered-response model to quantify misclassification rates in injury severity and adjust the bias in parameter estimates associated with misclassification. Misclassification was handled by treating the probability of observed injury severity as the product of the misclassification matrix and the latent injury severity variable—so-called true injury severity. The estimated misclassification matrix αi,j, the probability of classifying i to j, shows that 31.7 percent of possible injuries can be misclassified as no injury, and 29.8 percent of non-incapacitating injuries can be misclassified as possible injury.
Li, Kim, and Nitz (1999) used a logistic regression model that allows for misclassification errors in outcome variable to identify factors contributing to safety belt use overreported among crash-involved drivers and front seat passengers. Specifically, misclassification rates were represented by sensitivity (the probability of correctly classifying the safety belt use) and specificity (the probability of correctly classifying the absence of the safety belt), and both parameters were estimated along with other coefficients.
Several methods are identified as applicable to address the misreporting of motor vehicle crashes related to alcohol- and drug-impaired driving:
Traffic record linkage of crash and alternative data sources is primarily used to enrich the available data sets for safety analysis; however, linked data have been used to identify misreporting, especially underreporting. One of the most important issues in data linkage is how the databases are linked (e.g., deterministic, probabilistic method), as accuracy of linkage depends on the method.
When crash databases are not complete or a linkage method is not robust, the capture–recapture (or mark–recapture) method can be considered. This method uses linked data to compare law enforcement–reported records with an alternative data source, then estimates closer-to-true crashes using the overlapped samples of the two sources. Thus, underreporting or misreporting can be addressed, and unbiased analysis can be conducted.
Crash narratives can be used to recover missing or misclassified crashes. As part of the crash report, a crash narrative does not need a linkage but requires review of the text. Text mining can be employed to automate text data analysis. Thus, the technique can be used to capture missing or incomplete information in the narrative data related to driver impairment or distraction.
Lastly, researchers have developed an array of statistical models to estimate the underreporting rate of crashes and tackle the underreporting issue using only reported crash data. Although the models were developed for all types of crashes, they can be adapted to estimate the expected frequency of impaired and distracted driving crashes.