{"id":844,"date":"2026-09-26T12:42:32","date_gmt":"2026-09-26T12:42:32","guid":{"rendered":"https:\/\/perit.ai\/blogs\/?p=844"},"modified":"2026-09-26T12:42:32","modified_gmt":"2026-09-26T12:42:32","slug":"engineering-quality-in-ai-data-combining-human-expertise-with-automated-qa","status":"publish","type":"post","link":"https:\/\/perit.ai\/blogs\/engineering-quality-in-ai-data-combining-human-expertise-with-automated-qa\/","title":{"rendered":"Engineering Quality in AI Data: Combining Human Expertise with Automated QA"},"content":{"rendered":"<h2 style=\"margin-top: 24px; margin-bottom: 8px;\">Where Does Quality Break Down in AI Data?<\/h2>\n<div style=\"border-left: 3px solid #4A90E2; padding: 4px 0 4px 18px; margin: 25px 0; font-size: 16px; line-height: 1.7; color: #333;\">AI models depend heavily on the quality of the data used to train them. But as datasets grow to thousands or millions of images, videos, transcripts, and other data points, keeping every annotation consistent becomes more difficult. Most of the data may be correct, but even a small number of errors can affect the quality of the final dataset.<\/div>\n<p dir=\"auto\" data-start=\"964\" data-end=\"1238\">Some problems are easy to identify, such as a missing label or an incorrect value. Others are less obvious. Two annotators might look at the same image and choose different labels, or an unclear part of a transcript might be interpreted differently depending on the context.<\/p>\n<p dir=\"auto\" data-start=\"1240\" data-end=\"1287\">These problems generally fall into three areas:<\/p>\n<div style=\"display: flex; gap: 14px; margin: 28px 0;\">\n<div style=\"flex: 1; border: 1px solid #9fc7ee; border-radius: 6px; padding: 18px; background: linear-gradient(135deg, #f4f9ff 0%, #eaf4ff 100%); box-shadow: 0 2px 6px rgba(60,120,180,0.08);\">\n<p><strong style=\"font-size: 15px; color: #245a91;\">Inconsistency<\/strong><\/p>\n<p style=\"margin: 10px 0 0; font-size: 14px; line-height: 1.6; color: #333;\">Annotators may interpret the same guidelines differently.<\/p>\n<\/div>\n<div style=\"flex: 1; border: 1px solid #9fc7ee; border-radius: 6px; padding: 18px; background: linear-gradient(135deg, #f4f9ff 0%, #eaf4ff 100%); box-shadow: 0 2px 6px rgba(60,120,180,0.08);\">\n<p><strong style=\"font-size: 15px; color: #245a91;\">Annotation errors<\/strong><\/p>\n<p style=\"margin: 10px 0 0; font-size: 14px; line-height: 1.6; color: #333;\">Labels or other details may be incorrect or incomplete.<\/p>\n<\/div>\n<div style=\"flex: 1; border: 1px solid #9fc7ee; border-radius: 6px; padding: 18px; background: linear-gradient(135deg, #f4f9ff 0%, #eaf4ff 100%); box-shadow: 0 2px 6px rgba(60,120,180,0.08);\">\n<p><strong style=\"font-size: 15px; color: #245a91;\">Ambiguity<\/strong><\/p>\n<p style=\"margin: 10px 0 0; font-size: 14px; line-height: 1.6; color: #333;\">Some cases require context or domain knowledge before a decision can be made.<\/p>\n<\/div>\n<\/div>\n<p dir=\"auto\" data-start=\"1545\" data-end=\"1860\">When a dataset is small, these issues can often be found through manual review. As the volume increases, checking everything by hand becomes slower and harder to manage. This is where automated checks and human review can complement each other, with each handling the types of problems it is better suited to catch.<\/p>\n<h2 style=\"margin-top: 24px; margin-bottom: 8px;\">Why Small Errors Matter at Scale<\/h2>\n<p dir=\"auto\" data-start=\"1899\" data-end=\"2086\">The impact of an individual annotation error may seem small when looking at a single record. The situation changes when the same type of mistake appears repeatedly across a large dataset.<\/p>\n<p dir=\"auto\" data-start=\"2088\" data-end=\"2410\">For example, if an annotation guideline is interpreted differently by different reviewers, the resulting inconsistencies can spread across thousands of records. Similarly, a validation rule that is too broad may repeatedly flag valid annotations, while a rule that is too narrow may allow certain errors to pass unnoticed.<\/p>\n<div style=\"margin: 35px auto; max-width: 900px; text-align: center;\">\n<div style=\"display: flex; align-items: center; justify-content: center; gap: 10px; flex-wrap: wrap;\">\n<div style=\"width: 190px; min-height: 80px; display: flex; flex-direction: column; align-items: center; justify-content: center; padding: 14px; box-sizing: border-box; border: 1px solid #9fc7ee; border-radius: 6px; background: #f3f8ff;\"><strong style=\"font-size: 15px; color: #245a91;\">1. Single Error<\/strong><br \/>\n<span style=\"font-size: 13px; line-height: 1.4; color: #444; margin-top: 7px;\">One annotation is incorrect or inconsistent.<\/span><\/div>\n<div style=\"font-size: 24px; color: #4a90e2; flex-shrink: 0;\">\u2192<\/div>\n<div style=\"width: 190px; min-height: 80px; display: flex; flex-direction: column; align-items: center; justify-content: center; padding: 14px; box-sizing: border-box; border: 1px solid #9fc7ee; border-radius: 6px; background: #f3f8ff;\"><strong style=\"font-size: 15px; color: #245a91;\">2. Repeated Pattern<\/strong><br \/>\n<span style=\"font-size: 13px; line-height: 1.4; color: #444; margin-top: 7px;\">The same issue appears across multiple records.<\/span><\/div>\n<div style=\"font-size: 24px; color: #4a90e2; flex-shrink: 0;\">\u2192<\/div>\n<div style=\"width: 190px; min-height: 80px; display: flex; flex-direction: column; align-items: center; justify-content: center; padding: 14px; box-sizing: border-box; border: 1px solid #9fc7ee; border-radius: 6px; background: #eaf4ff;\"><strong style=\"font-size: 15px; color: #245a91;\">3. Dataset-Level Impact<\/strong><br \/>\n<span style=\"font-size: 13px; line-height: 1.4; color: #444; margin-top: 7px;\">A recurring error can affect thousands of annotations.<\/span><\/div>\n<\/div>\n<\/div>\n<p dir=\"auto\" data-start=\"2412\" data-end=\"2593\">This makes dataset quality more than a matter of finding individual mistakes. The process also needs to identify <strong data-start=\"2525\" data-end=\"2555\">patterns in those mistakes<\/strong> and determine why they are happening.<\/p>\n<p dir=\"auto\" data-start=\"2595\" data-end=\"2655\">A useful QA process therefore needs to answer two questions:<\/p>\n<p dir=\"auto\" data-start=\"2657\" data-end=\"2757\"><strong data-start=\"2657\" data-end=\"2757\">Which annotations need attention, and what is causing those issues to appear in the first place?<\/strong><\/p>\n<p dir=\"auto\" data-start=\"2759\" data-end=\"2958\">Automated checks can help identify large numbers of potentially problematic records. Human reviewers can then examine the cases where context, interpretation, or project-specific knowledge is needed.<\/p>\n<h2 style=\"margin-top: 24px; margin-bottom: 8px;\">Automated QA at Scale<\/h2>\n<p dir=\"auto\" data-start=\"2986\" data-end=\"3205\">When AI datasets grow to thousands or millions of records, checking every annotation manually is not practical. Automated QA provides a first layer of validation by applying the same rules across a large volume of data.<\/p>\n<p dir=\"auto\" data-start=\"3207\" data-end=\"3445\">For example, an automated system can flag a missing label, an invalid category, a duplicate annotation, or a bounding box that extends beyond an image. These are problems that can be described through clear rules and checked consistently.<\/p>\n<div style=\"margin: 35px 0; text-align: center;\">\n<div style=\"display: flex; align-items: center; justify-content: center; gap: 10px; flex-wrap: wrap;\">\n<div style=\"border: 1px solid #9fc7ee; border-radius: 6px; padding: 14px 20px; background: #f3f8ff; min-width: 110px;\"><strong style=\"color: #245a91;\">Dataset<\/strong><\/div>\n<div style=\"font-size: 22px; color: #4a90e2;\">\u2192<\/div>\n<div style=\"border: 1px solid #9fc7ee; border-radius: 6px; padding: 14px 20px; background: #edf6ff; min-width: 150px;\"><strong style=\"color: #245a91;\">Automated<br \/>\nValidation<\/strong><\/div>\n<div style=\"font-size: 22px; color: #4a90e2;\">\u2192<\/div>\n<div style=\"border: 2px solid #4A90E2; border-radius: 50%; width: 100px; height: 100px; display: flex; align-items: center; justify-content: center; background: #f4f9ff; padding: 8px; box-sizing: border-box;\"><strong style=\"font-size: 14px; color: #245a91;\">Issues<br \/>\nDetected?<\/strong><\/div>\n<\/div>\n<div style=\"display: flex; justify-content: center; gap: 70px; margin-top: 10px; font-size: 13px;\">\n<div style=\"color: #4a90e2;\"><strong>No \u2193<\/strong><\/div>\n<div style=\"color: #4a90e2;\"><strong>Yes \u2192<\/strong><\/div>\n<\/div>\n<div style=\"display: flex; justify-content: center; align-items: center; gap: 10px; margin-top: 8px; flex-wrap: wrap;\">\n<div style=\"border: 1px solid #b7dfd5; border-radius: 6px; padding: 14px 20px; background: #f1faf7; min-width: 150px;\"><strong style=\"color: #26705d;\">Continue<br \/>\nValidation<\/strong><\/div>\n<div style=\"font-size: 22px; color: #4a90e2;\">\u2190<\/div>\n<div style=\"border: 1px solid #9fc7ee; border-radius: 6px; padding: 14px 20px; background: #f3f8ff; min-width: 130px;\"><strong style=\"color: #245a91;\">Human<br \/>\nReview<\/strong><\/div>\n<div style=\"font-size: 22px; color: #4a90e2;\">\u2192<\/div>\n<div style=\"border: 1px solid #b7dfd5; border-radius: 6px; padding: 14px 20px; background: #f1faf7; min-width: 150px;\"><strong style=\"color: #26705d;\">Validated<br \/>\nDataset<\/strong><\/div>\n<\/div>\n<\/div>\n<p dir=\"auto\" data-start=\"3447\" data-end=\"3686\">Within <a class=\"decorated-link\" href=\"https:\/\/perit.ai\/services\/data-annotation?utm_source=chatgpt.com\" target=\"_new\" rel=\"noopener\" data-start=\"3454\" data-end=\"3535\"><strong data-start=\"3455\" data-end=\"3491\">Perit\u2019s data annotation workflow<\/strong><\/a>, automation can help identify cases that need a closer look. Instead of replacing human reviewers, it can narrow down where their attention is needed.<\/p>\n<p dir=\"auto\" data-start=\"3688\" data-end=\"3972\">This creates a practical division of work. <strong data-start=\"3731\" data-end=\"3843\">Automation handles repeatable checks, while human reviewers focus on cases that require context or judgment.<\/strong> The approach becomes particularly useful when the same quality standards need to be maintained across large and varied datasets.<\/p>\n<h2 style=\"margin-top: 24px; margin-bottom: 8px;\">The Role of Human Review<\/h2>\n<p dir=\"auto\" data-start=\"4003\" data-end=\"4255\">Automation is effective when an issue can be defined by a clear rule. However, not every annotation can be judged that way. Some cases require a reviewer to understand the surrounding context before deciding whether an annotation is actually incorrect.<\/p>\n<p dir=\"auto\" data-start=\"4257\" data-end=\"4495\">Consider a dataset containing <strong data-start=\"4287\" data-end=\"4308\">1,000 annotations<\/strong>. An automated QA system flags 80 of them as potential issues. A human reviewer examines those 80 cases and finds that 60 are genuine errors, while the remaining 20 are valid annotations.<\/p>\n<p dir=\"auto\" data-start=\"4497\" data-end=\"4584\">We can use <strong data-start=\"4508\" data-end=\"4521\">precision<\/strong> to measure how many of the flagged cases were actually errors.<\/p>\n<div style=\"text-align: center; margin: 25px 0;\">\n<table style=\"margin: 0 auto; border-collapse: collapse; border: none; width: auto; font-size: 18px; line-height: 1.2;\">\n<tbody>\n<tr>\n<td style=\"border: none; padding: 0 10px 0 0; white-space: nowrap; vertical-align: middle;\" rowspan=\"2\"><strong>Precision =<\/strong><\/td>\n<td style=\"border: none; border-bottom: 1px solid #333; padding: 0 14px 6px; text-align: center; white-space: nowrap;\"><strong>True Positives<\/strong><\/td>\n<\/tr>\n<tr>\n<td style=\"border: none; padding: 6px 14px 0; text-align: center; white-space: nowrap;\"><strong>True Positives + False Positives<\/strong><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p dir=\"auto\" data-start=\"4655\" data-end=\"4671\">In this example:<\/p>\n<div style=\"text-align: center; margin: 18px 0; font-size: 17px;\"><strong>Precision = 60 \/ (60 + 20) = 75%<\/strong><\/div>\n<p dir=\"auto\" data-start=\"4711\" data-end=\"4815\">This means that <strong data-start=\"4727\" data-end=\"4814\">75% of the annotations flagged by the automated system actually required correction<\/strong>.<\/p>\n<p dir=\"auto\" data-start=\"4817\" data-end=\"5051\">The example also shows why human review remains important. An automated system may identify a pattern, but it cannot always determine whether an unusual case is genuinely wrong or simply an exception that the project guidelines allow.<\/p>\n<p dir=\"auto\" data-start=\"5053\" data-end=\"5259\">Human review can also expose recurring problems. If reviewers repeatedly find the same type of annotation error, the issue may point to unclear instructions or missing examples in the annotation guidelines.<\/p>\n<h2 style=\"margin-top: 24px; margin-bottom: 8px;\">Measuring Data Quality<\/h2>\n<p dir=\"auto\" data-start=\"5288\" data-end=\"5482\">Finding errors is only one part of QA. For large datasets, teams also need a way to understand how frequently those errors occur and whether the annotation process is being applied consistently.<\/p>\n<p dir=\"auto\" data-start=\"5484\" data-end=\"5597\">One simple measure is the <strong data-start=\"5510\" data-end=\"5524\">error rate<\/strong>, which represents the proportion of annotations that require correction.<\/p>\n<div style=\"overflow-x: auto; margin: 28px 0;\">\n<table style=\"width: 100%; border-collapse: collapse; font-size: 14px; line-height: 1.5;\">\n<tbody>\n<tr>\n<th style=\"background: #eaf4ff; color: #245a91; border: 1px solid #b9d5ee; padding: 12px; text-align: left;\">Metric<\/th>\n<th style=\"background: #eaf4ff; color: #245a91; border: 1px solid #b9d5ee; padding: 12px; text-align: left;\">What it tells us<\/th>\n<th style=\"background: #eaf4ff; color: #245a91; border: 1px solid #b9d5ee; padding: 12px; text-align: left;\">Why it matters<\/th>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #d5dfe8; padding: 12px;\"><strong>Error rate<\/strong><\/td>\n<td style=\"border: 1px solid #d5dfe8; padding: 12px;\">How much of the dataset requires correction.<\/td>\n<td style=\"border: 1px solid #d5dfe8; padding: 12px;\">Shows the overall level of incorrect annotations.<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #d5dfe8; padding: 12px;\"><strong>Precision<\/strong><\/td>\n<td style=\"border: 1px solid #d5dfe8; padding: 12px;\">How many flagged cases are actual errors.<\/td>\n<td style=\"border: 1px solid #d5dfe8; padding: 12px;\">Shows how useful automated flags are.<\/td>\n<\/tr>\n<tr>\n<td style=\"border: 1px solid #d5dfe8; padding: 12px;\"><strong>Inter-annotator agreement<\/strong><\/td>\n<td style=\"border: 1px solid #d5dfe8; padding: 12px;\">How consistently reviewers apply the guidelines.<\/td>\n<td style=\"border: 1px solid #d5dfe8; padding: 12px;\">Helps identify unclear or inconsistently applied guidelines.<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<div style=\"text-align: center; margin: 28px 0;\">\n<table style=\"margin: 0 auto; border-collapse: collapse; border: none; width: auto; font-size: 18px; line-height: 1.2;\">\n<tbody>\n<tr>\n<td style=\"border: none; padding: 0 10px 0 0; white-space: nowrap; vertical-align: middle;\" rowspan=\"2\"><strong>Error Rate =<\/strong><\/td>\n<td style=\"border: none; border-bottom: 1px solid #333; padding: 0 14px 6px; text-align: center; white-space: nowrap;\"><strong>Incorrect Annotations<\/strong><\/td>\n<td style=\"border: none; padding: 0 0 0 10px; white-space: nowrap; vertical-align: middle;\" rowspan=\"2\"><strong>\u00d7 100<\/strong><\/td>\n<\/tr>\n<tr>\n<td style=\"border: none; padding: 6px 14px 0; text-align: center; white-space: nowrap;\"><strong>Total Annotations<\/strong><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p dir=\"auto\" data-start=\"5665\" data-end=\"5755\">For example, if a batch contains <strong data-start=\"5698\" data-end=\"5719\">2,000 annotations<\/strong> and 50 are found to contain errors:<\/p>\n<div><\/div>\n<div style=\"text-align: center; margin: 18px 0; font-size: 17px;\"><strong>Error Rate = (50 \/ 2,000) \u00d7 100 = 2.5%<\/strong><\/div>\n<div><\/div>\n<div style=\"margin: 30px 0; border: 1px solid #d5dfe8; border-radius: 6px; padding: 22px 24px; background: #fafcff;\">\n<div style=\"text-align: center; margin-bottom: 18px;\"><strong style=\"font-size: 16px; color: #245a91;\"><br \/>\nIllustrative Example: Annotation Quality in a Batch<br \/>\n<\/strong><\/div>\n<div style=\"display: flex; width: 100%; height: 42px; border-radius: 5px; overflow: hidden;\">\n<div style=\"width: 97.5%; background: #dceeff; display: flex; align-items: center; justify-content: center; color: #245a91; font-size: 14px; font-weight: bold;\">97.5% \u2014 No identified errors<\/div>\n<div style=\"width: 2.5%; background: #4A90E2; min-width: 55px; display: flex; align-items: center; justify-content: center; color: white; font-size: 12px; font-weight: bold;\">2.5%<\/div>\n<\/div>\n<div style=\"display: flex; justify-content: space-between; margin-top: 12px; font-size: 13px; color: #555;\">1,950 annotations<br \/>\n50 annotations<\/div>\n<\/div>\n<p dir=\"auto\" data-start=\"5799\" data-end=\"5992\">This gives a simple view of how much incorrect work was identified in the batch. However, error rate alone does not explain whether the annotation guidelines are being interpreted consistently.<\/p>\n<p dir=\"auto\" data-start=\"5994\" data-end=\"6037\">That is where other measures can be useful.<\/p>\n<h2 style=\"margin-top: 24px; margin-bottom: 8px;\">Turning QA Results Into Measurable Signals<\/h2>\n<p dir=\"auto\" data-start=\"6086\" data-end=\"6357\"><strong data-start=\"6086\" data-end=\"6115\">Inter-annotator agreement<\/strong> looks at how often independent annotators reach the same decision when working on the same data. When reviewers consistently reach different conclusions, it can indicate that a guideline is unclear or that certain cases need better examples.<\/p>\n<p dir=\"auto\" data-start=\"6359\" data-end=\"6508\"><strong data-start=\"6359\" data-end=\"6372\">Precision<\/strong> provides another perspective by showing how many of the cases flagged by an automated check turn out to be genuine errors after review.<\/p>\n<div style=\"display: flex; gap: 16px; margin: 30px 0; flex-wrap: wrap;\">\n<div style=\"flex: 1; min-width: 180px; padding: 18px; border-top: 3px solid #4A90E2; background: #f5f9fe;\">\n<p><strong style=\"color: #245a91; font-size: 15px;\">Error Rate<\/strong><\/p>\n<p style=\"margin: 8px 0 0; font-size: 14px; line-height: 1.5;\">How much incorrect work was found in the dataset.<\/p>\n<\/div>\n<div style=\"flex: 1; min-width: 180px; padding: 18px; border-top: 3px solid #4A90E2; background: #f5f9fe;\">\n<p><strong style=\"color: #245a91; font-size: 15px;\">Precision<\/strong><\/p>\n<p style=\"margin: 8px 0 0; font-size: 14px; line-height: 1.5;\">How many flagged cases were confirmed as genuine errors.<\/p>\n<\/div>\n<div style=\"flex: 1; min-width: 180px; padding: 18px; border-top: 3px solid #4A90E2; background: #f5f9fe;\">\n<p><strong style=\"color: #245a91; font-size: 15px;\">Inter-annotator Agreement<\/strong><\/p>\n<p style=\"margin: 8px 0 0; font-size: 14px; line-height: 1.5;\">How consistently reviewers apply the same guidelines.<\/p>\n<\/div>\n<\/div>\n<p dir=\"auto\" data-start=\"6510\" data-end=\"6791\">These measurements answer different questions. Error rate helps show <strong data-start=\"6579\" data-end=\"6616\">how much incorrect work was found<\/strong>, precision helps show <strong data-start=\"6639\" data-end=\"6686\">how useful an automated flagging process is<\/strong>, and inter-annotator agreement helps show <strong data-start=\"6729\" data-end=\"6790\">how consistently the annotation standard is being applied<\/strong>.<\/p>\n<p dir=\"auto\" data-start=\"6793\" data-end=\"7018\">For projects handled through <a class=\"decorated-link\" href=\"https:\/\/perit.ai\/services\/data-annotation?utm_source=chatgpt.com\" target=\"_new\" rel=\"noopener\" data-start=\"6822\" data-end=\"6903\"><strong data-start=\"6823\" data-end=\"6859\">Perit\u2019s data annotation workflow<\/strong><\/a>, these measurements can be considered alongside sampling and human review to understand how a batch is performing.<\/p>\n<p dir=\"auto\" data-start=\"7020\" data-end=\"7241\">The important point is that no single metric can describe the complete quality of a dataset. Looking at several signals together provides a more useful picture of where the process is working and where it needs attention.<\/p>\n<h2 style=\"margin-top: 24px; margin-bottom: 8px;\">A Hybrid QA Workflow<\/h2>\n<p dir=\"auto\" data-start=\"7268\" data-end=\"7369\">The real value of combining automation with human review comes from how the two stages work together.<\/p>\n<p dir=\"auto\" data-start=\"7371\" data-end=\"7735\">A typical workflow can start with annotated data passing through automated validation. Records that meet the defined checks can continue through the process, while potential issues are flagged for closer inspection. Human reviewers can then determine whether the flagged cases are genuine errors, valid exceptions, or cases where the guideline needs clarification.<\/p>\n<p dir=\"auto\" data-start=\"7737\" data-end=\"7772\">The workflow can be represented as:<\/p>\n<div style=\"margin: 35px 0; overflow-x: auto;\">\n<div style=\"display: flex; align-items: center; justify-content: center; gap: 8px; min-width: 850px;\">\n<div style=\"padding: 13px 14px; border: 1px solid #9fc7ee; border-radius: 5px; background: #f4f9ff; text-align: center; min-width: 80px;\"><strong style=\"font-size: 13px; color: #245a91;\">Data<\/strong><\/div>\n<p><span style=\"font-size: 20px; color: #4a90e2;\">\u2192<\/span><\/p>\n<div style=\"padding: 13px 14px; border: 1px solid #9fc7ee; border-radius: 5px; background: #f4f9ff; text-align: center; min-width: 90px;\"><strong style=\"font-size: 13px; color: #245a91;\">Annotation<\/strong><\/div>\n<p><span style=\"font-size: 20px; color: #4a90e2;\">\u2192<\/span><\/p>\n<div style=\"padding: 13px 14px; border: 1px solid #9fc7ee; border-radius: 5px; background: #eaf4ff; text-align: center; min-width: 125px;\"><strong style=\"font-size: 13px; color: #245a91;\">Automated<br \/>\nValidation<\/strong><\/div>\n<p><span style=\"font-size: 20px; color: #4a90e2;\">\u2192<\/span><\/p>\n<div style=\"padding: 13px 14px; border: 1px solid #9fc7ee; border-radius: 5px; background: #eaf4ff; text-align: center; min-width: 105px;\"><strong style=\"font-size: 13px; color: #245a91;\">Flagged<br \/>\nCases<\/strong><\/div>\n<p><span style=\"font-size: 20px; color: #4a90e2;\">\u2192<\/span><\/p>\n<div style=\"padding: 13px 14px; border: 1px solid #9fc7ee; border-radius: 5px; background: #f4f9ff; text-align: center; min-width: 100px;\"><strong style=\"font-size: 13px; color: #245a91;\">Human<br \/>\nReview<\/strong><\/div>\n<p><span style=\"font-size: 20px; color: #4a90e2;\">\u2192<\/span><\/p>\n<div style=\"padding: 13px 14px; border: 1px solid #9fc7ee; border-radius: 5px; background: #f4f9ff; text-align: center; min-width: 90px;\"><strong style=\"font-size: 13px; color: #245a91;\">Correction<\/strong><\/div>\n<p><span style=\"font-size: 20px; color: #4a90e2;\">\u2192<\/span><\/p>\n<div style=\"padding: 13px 14px; border: 1px solid #9fc7ee; border-radius: 5px; background: #eaf4ff; text-align: center; min-width: 90px;\"><strong style=\"font-size: 13px; color: #245a91;\">Final QA<\/strong><\/div>\n<p><span style=\"font-size: 20px; color: #4a90e2;\">\u2192<\/span><\/p>\n<div style=\"padding: 13px 14px; border: 1px solid #a8d8c9; border-radius: 5px; background: #f1faf7; text-align: center; min-width: 105px;\"><strong style=\"font-size: 13px; color: #26705d;\">Validated<br \/>\nDataset<\/strong><\/div>\n<\/div>\n<\/div>\n<div style=\"text-align: center; margin: -10px 0 30px; font-size: 13px; color: #666;\">QA findings can be used to refine guidelines and validation rules<\/div>\n<p dir=\"auto\" data-start=\"7895\" data-end=\"8198\">At <a class=\"decorated-link\" href=\"https:\/\/perit.ai\/services\/data-annotation?utm_source=chatgpt.com\" target=\"_new\" rel=\"noopener\" data-start=\"7898\" data-end=\"7952\"><strong data-start=\"7899\" data-end=\"7908\">Perit<\/strong><\/a>, this type of quality-focused process can be supported through structured annotation, calibration, sampling, and review. Its <a class=\"decorated-link\" href=\"https:\/\/foundry.zyka.ai\/\" target=\"_new\" rel=\"noopener\" data-start=\"8078\" data-end=\"8117\"><strong data-start=\"8079\" data-end=\"8090\">Foundry<\/strong><\/a> workspace provides the environment for annotation and related review activities.<\/p>\n<p dir=\"auto\" data-start=\"8200\" data-end=\"8489\">The workflow does not have to end when an annotation is corrected. Findings from human review can also be used to improve the next stage of the process. A repeated error may require a clearer guideline, while a validation rule that generates too many unnecessary flags may need refinement.<\/p>\n<p dir=\"auto\" data-start=\"8491\" data-end=\"8581\">This turns QA from a final inspection into an ongoing part of the data production process.<\/p>\n<h2 style=\"margin-top: 24px; margin-bottom: 8px;\">Learning From QA Results<\/h2>\n<p dir=\"auto\" data-start=\"8612\" data-end=\"8744\">A useful QA system should not only identify what went wrong. It should also help prevent the same problem from appearing repeatedly.<\/p>\n<p dir=\"auto\" data-start=\"8746\" data-end=\"8979\">Suppose reviewers keep correcting the same type of label. The issue may not necessarily be with the individual annotators. The guideline might need a clearer definition, an additional example, or a better explanation of an edge case.<\/p>\n<p dir=\"auto\" data-start=\"8981\" data-end=\"9200\">The same applies to automated checks. If a rule repeatedly flags valid annotations, the rule may need to be adjusted. If a particular type of error is repeatedly missed, a new validation check may need to be introduced.<\/p>\n<p dir=\"auto\" data-start=\"9202\" data-end=\"9231\">This creates a feedback loop:<\/p>\n<div style=\"margin: 32px auto 25px; max-width: 1000px; text-align: center;\">\n<div style=\"display: flex; align-items: center; justify-content: center; gap: 8px; flex-wrap: nowrap;\">\n<div style=\"width: 150px; min-height: 64px; display: flex; align-items: center; justify-content: center; padding: 12px; box-sizing: border-box; border: 1px solid #9fc7ee; border-radius: 6px; background: #f3f8ff; color: #245a91; font-size: 15px; font-weight: bold;\">QA Result<\/div>\n<div style=\"font-size: 21px; color: #4a90e2; flex-shrink: 0;\">\u2192<\/div>\n<div style=\"width: 150px; min-height: 64px; display: flex; align-items: center; justify-content: center; padding: 12px; box-sizing: border-box; border: 1px solid #9fc7ee; border-radius: 6px; background: #f3f8ff; color: #245a91; font-size: 15px; font-weight: bold;\">Error Analysis<\/div>\n<div style=\"font-size: 21px; color: #4a90e2; flex-shrink: 0;\">\u2192<\/div>\n<div style=\"width: 165px; min-height: 64px; display: flex; align-items: center; justify-content: center; padding: 12px; box-sizing: border-box; border: 1px solid #9fc7ee; border-radius: 6px; background: #eaf4ff; color: #245a91; font-size: 15px; font-weight: bold; line-height: 1.35;\">Guideline \/ Rule<br \/>\nUpdate<\/div>\n<div style=\"font-size: 21px; color: #4a90e2; flex-shrink: 0;\">\u2192<\/div>\n<div style=\"width: 150px; min-height: 64px; display: flex; align-items: center; justify-content: center; padding: 12px; box-sizing: border-box; border: 1px solid #9fc7ee; border-radius: 6px; background: #f3f8ff; color: #245a91; font-size: 15px; font-weight: bold;\">Re-annotation<\/div>\n<div style=\"font-size: 21px; color: #4a90e2; flex-shrink: 0;\">\u2192<\/div>\n<div style=\"width: 150px; min-height: 64px; display: flex; align-items: center; justify-content: center; padding: 12px; box-sizing: border-box; border: 1px solid #a9d8cc; border-radius: 6px; background: #f0faf7; color: #246b5c; font-size: 15px; font-weight: bold;\">New QA Check<\/div>\n<\/div>\n<div style=\"margin-top: 14px; color: #4a90e2; font-size: 20px; line-height: 1;\"><\/div>\n<div style=\"margin-top: 3px; font-size: 13px; color: #666;\">Continuous improvement<\/div>\n<\/div>\n<p dir=\"auto\" data-start=\"9323\" data-end=\"9582\">For <a class=\"decorated-link\" href=\"https:\/\/perit.ai\/services\/data-annotation?utm_source=chatgpt.com\" target=\"_new\" rel=\"noopener\" data-start=\"9327\" data-end=\"9381\"><strong data-start=\"9328\" data-end=\"9337\">Perit<\/strong><\/a>, calibration and review are important parts of maintaining consistent annotation standards. QA findings can provide additional information about where those standards or processes may need refinement.<\/p>\n<p dir=\"auto\" data-start=\"9584\" data-end=\"9784\">The goal is not simply to correct more errors in the current batch. It is to <strong data-start=\"9661\" data-end=\"9729\">reduce the chance of the same errors appearing in future batches<\/strong>, making the overall process more consistent over time.<\/p>\n<h2 style=\"margin-top: 24px; margin-bottom: 8px;\">Perit\u2019s Approach to Reliable AI Data<\/h2>\n<p dir=\"auto\" data-start=\"9827\" data-end=\"9967\">For complex AI projects, quality assurance is most useful when it is built into the data workflow rather than treated as a final inspection.<\/p>\n<p dir=\"auto\" data-start=\"9969\" data-end=\"10259\">Automated validation can handle repeatable checks across large datasets, while human reviewers can focus on cases where context and judgment matter. Measurements such as error rate, precision, and inter-annotator agreement can then provide a way to understand how the process is performing.<\/p>\n<p dir=\"auto\" data-start=\"10261\" data-end=\"10570\">Perit brings these elements together through its <a class=\"decorated-link\" href=\"https:\/\/perit.ai\/services\/data-annotation?utm_source=chatgpt.com\" target=\"_new\" rel=\"noopener\" data-start=\"10310\" data-end=\"10383\"><strong data-start=\"10311\" data-end=\"10339\">data annotation workflow<\/strong><\/a>, supported by calibration, sampling, review, and quality measurement. Its <a class=\"decorated-link\" href=\"https:\/\/foundry.zyka.ai\/\" target=\"_new\" rel=\"noopener\" data-start=\"10458\" data-end=\"10497\"><strong data-start=\"10459\" data-end=\"10470\">Foundry<\/strong><\/a> workspace supports the annotation process and related review activities.<\/p>\n<p dir=\"auto\" data-start=\"10572\" data-end=\"10609\">The approach can be summed up simply:<\/p>\n<div style=\"border-left: 3px solid #4A90E2; padding: 2px 0 2px 18px; margin: 22px 0; font-size: 16px; line-height: 1.7; color: #333;\"><strong>AI models depend heavily on the quality of the data used to train them. But as datasets grow to thousands or millions of images, videos, transcripts, and other data points, keeping every annotation consistent becomes more difficult.<\/strong><\/div>\n<p dir=\"auto\" data-start=\"10752\" data-end=\"10941\">For teams working with large and complex AI datasets, combining automated validation with human expertise provides a practical way to make quality more consistent, measurable, and scalable.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Where Does Quality Break Down in AI Data? AI models depend heavily on the quality of the data used to train them. But as datasets grow to thousands or millions\u2026<\/p>\n","protected":false},"author":1,"featured_media":871,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[22,19],"tags":[],"class_list":["post-844","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-artificial-intelligence","category-research"],"_links":{"self":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/844","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/comments?post=844"}],"version-history":[{"count":5,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/844\/revisions"}],"predecessor-version":[{"id":872,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/posts\/844\/revisions\/872"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/media\/871"}],"wp:attachment":[{"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/media?parent=844"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/categories?post=844"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/perit.ai\/blogs\/wp-json\/wp\/v2\/tags?post=844"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}