On this page
- Abstract
- 1 Introduction
- 2 Related work
- 2.1 Manufacturing test sequences
- 2.2 Data-driven test picking models
- 2.3 Multi-armed bandit algorithms
- 3 Task and model definition
- 4 Experiments
- 4.1 Experiment design
- 4.2 Results
- 5 Discussion
- Data availability statement
- Author contributions
- Funding
- Conflict of interest
- Generative AI statement
- Publisher’s note
- Article notes
- References
Abstract
Introduction: Manufacturing test flows in high-volume electronics production are typically fixed during product development and executed unchanged on every unit, even as failure patterns and process conditions evolve. This protects quality, but it also imposes unnecessary test cost, while existing data-driven methods mostly optimize static test subsets and neither adapt online to changing defect distributions nor explicitly control escape risk.Methods: In this study, we present an adaptive test-selection framework that combines offline minimum-cost diagnostic subset construction using greedy set cover with an online Thompson-sampling multi-armed bandit that switches between full and reduced test plans using a rolling process-stability signal. We evaluate the framework on two printed circuit board assembly stages—Functional Circuit Test and End-of-Line test—covering 28,000 board runs.Results: Offline analysis identified zero-escape reduced plans that cut test time by 18.78% in Functional Circuit Test and 91.57% in End-of-Line testing. Under temporal validation with real concept drift, static reduction produced 110 escaped defects in Functional Circuit Test and 8 in End-of-Line, whereas the adaptive policy reduced escapes to zero by reverting to fuller coverage when instability emerged in practice.Discussion: These results show that online learning can preserve manufacturing quality while reducing test burden, offering a practical route to adaptive test planning across production domains, and offering both economic and logistics improvement for companies.
1 Introduction
Ensuring that manufactured units conform to specified quality requirements before release is a fundamental objective of industrial production systems (Pharmaceutical Inspection Co-operation Scheme, 2009; Duarte et al., 2025; Grznár et al., 2025). When defective products escape into the field, the resulting consequences extend beyond the direct costs of repair, replacement, and rework to include reputational damage, customer dissatisfaction, and loss of future business (Walston et al., 2025; Tong et al., 2023; Bhowmick and Seetharaman, 2023). In high-volume manufacturing environments, these effects are amplified, making systematic quality assurance a critical component of operational performance (Charan Kantumchu et al., 2024; Lervåg Synnes and Welo, 2016). In this context, testing is one of the central mechanisms through which this assurance is achieved (Voorakaranam et al., 2002; Milor and Sangiovanni-Vincentelli, 1994). Across modern production lines, products are subjected to structured sequences of diagnostic and functional checks, including electrical measurements, communication verification, firmware validation, and system-level performance evaluation, with the aim of identifying defective units before shipment (Quigley, 2024; Kim et al., 2010; Chiara Magnanini et al., 2024).
Despite its importance, the design and execution of manufacturing test sequences remain challenging (Tahera et al., 2019; Chowdhury and Nuruzzaman, 2023; Fan et al., 2025). In most industrial settings, test flows are defined during product development on the basis of anticipated quality requirements and are then deployed as fixed procedures throughout production (Lackner et al., 2014; Tahera et al., 2012). Such staged verification strategies are widely used across domains. In semiconductor manufacturing, for example, early stages such as wafer sort assess die-level electrical integrity, whereas later stages such as package test verify the functionality of the packaged component before downstream processing (Iaria et al., 2025). Similarly, in biomedical additive manufacturing, early inspections evaluate dimensional accuracy and assembly integrity, while later stages address mechanical performance, sterilization, and regulatory compliance (Cecchitelli et al., 2026). Although this fixed and comprehensive approach supports consistency and traceability, it does not account for the inherently dynamic nature of manufacturing environments, in which process variability, material changes, equipment drift, and evolving defect modes may alter the diagnostic value of individual test steps over time (TRAN, 2021; Yang et al., 2025). Consequently, test steps that were once highly informative may become redundant, while previously uninformative steps may gain diagnostic relevance as production conditions change (Wong and Perumal, 2025). At the same time, contemporary manufacturing systems continuously generate large volumes of process and test data that could support more adaptive decision-making (Cerquitelli et al., 2021).
In parallel with these industrial developments, an increasing body of literature has examined the use of artificial intelligence (AI) and data-driven methods to improve manufacturing test efficiency (Jose Plathottam et al., 2023; Ghelani, 2024a; Ghelani, 2024b). Studies have shown that automation and learning-based approaches can improve repeatability and reduce operational cost in electronics testing environments (Anusuya et al., 2024; Vijay Agrawal et al., 2024). Other studies have investigated the selection of diagnostically informative test items from historical data. For example, Anusuya et al. (Anusuya et al., 2024) and Agrawal et al. (Vijay Agrawal et al., 2024) demonstrated that AI-driven automation improves repeatability and reduces operational cost in electronics manufacturing test processes, yet neither addresses which test steps should be executed in the first place. Pan et al. (Pan et al., 2025) proposed an ensemble-learning framework that filters test items according to diagnostic contribution and reported a 34% reduction in executed test items. However, the selected subset remains fixed after training, and the method does not address changes in defect distributions during production.
Namely, meaningful reductions in testing effort are generally achieved under the assumption of stable operating conditions (Kim et al., 2022). The problem of dynamically switching between full and reduced test plans during production, reverting to complete coverage when process conditions become unstable, and resuming reduced testing when stability is restored, remains largely unresolved. Equally important, prior approaches rarely quantify the defect escape risk associated with reduced testing or provide a mechanism for controlling that risk online. This limitation is particularly significant in manufacturing settings, where defect patterns may shift unexpectedly and where an initially safe reduced test plan may become insufficient when new failure modes emerge.
To this end, in this study, we address this gap by investigating whether historical production data can be used not only to identify a minimum-cost subset of diagnostically relevant test steps, but also to support an adaptive online policy that selects between full and reduced test plans on a per-unit basis. Specifically, we examine whether complete defect coverage can be preserved with a reduced subset derived from historical data, how the trade-off between test-time reduction and tolerated escape risk can be characterized, and whether an adaptive data-driven strategy can maintain quality performance under changing production conditions without offline retraining or manual reconfiguration. Formally, we propose an adaptive test-selection framework that combines offline subset optimization with online sequential decision-making. In the offline stage, a greedy set cover procedure is used to identify a minimum-cost diagnostic subset from historical production records (Adamo et al., 2023). In the online stage, a Thompson Sampling-based multi-armed bandit (MAB) model (Russo et al., 2018; Loecher, 2021) dynamically selects between the reduced and full test plans for each incoming unit using real-time production evidence. Notably, previous studies demonstrated the usefulness of MAB-based models for test reduction but focused on software-testing only (Antonio do Prado Lima and Vergilio, 2020). This two-stage formulation is intended to reduce testing effort during stable operating periods while preserving full coverage when instability is detected. The proposed model is designed for production environments in which test steps are executed sequentially, outcomes are measurable, and quality performance can be assessed in terms of defect detection.
In order to evaluate the proposed model’s performance, we explore its performance on Printed Circuit Board Assembly (PCBA) manufacturing conducted in collaboration with a large-scale industrial electronics manufacturing company. PCBA is a representative and demanding application domain characterized by high production volume, strict quality requirements, and multi-stage diagnostic testing. We include two sequential stages, Functional Circuit Test (FCT) and End-of-Line (EOL) testing, comprising 172 and 55 individual test steps, respectively. Historical data from more than 28,000 board runs are used to analyze defect patterns and construct reduced test plans, while temporally separated validation data are used to assess performance under real concept drift conditions.
The remainder of this paper is structured as follows. Section 2 briefly introduces manufacturing test optimization, data-driven test selection methods, and MAB algorithms for sequential decision-making. Section 3 formalizes the problem and outlines the proposed model. Section 4 describes the experimental setup, followed by the obtained results. Finally, Section 5 discusses the findings, limitations, and directions for future work.
2 Related work
In this section, we cover three core components of the proposed framework. First, the structure and economics of manufacturing test sequences are presented, establishing the context in which test optimization operates. Next, data-driven approaches to test subset selection and adaptive testing are outlined, with their limitations. Finally, we briefly present the MAB algorithm and its applications to sequential decision problems in manufacturing and quality control contexts, providing the theoretical foundation for the proposed model.
2.1 Manufacturing test sequences
Testing is a critical stage in the manufacturing process, ensuring that products meet design specifications and are free from quality defects. The problem of determining which tests should be executed and in what sequence, while minimizing total test time without compromising defect detection coverage, is a fundamental challenge in high-volume production environments (Boumen et al., 2008; Wang and Yang, 2025; Yilmaz et al., 2010). For instance, Scheffler et al. (Scheffler et al., 2004) demonstrate that the relationship between test effort and quality exhibits diminishing returns, such that additional testing yields progressively smaller improvements in defect reduction as fault coverage approaches saturation. In a similar manner, Walston et al. (Walston et al., 2025) situate quality-related costs within the broader cost of quality framework and highlight that poor quality can impose financial consequences beyond direct rework and scrap costs, including lost sales, lost profits, reputational damage, and loss of repeat business.
An important implication of this cost-coverage trade-off is that the diagnostic contribution of individual test steps becomes highly relevant (Pan et al., 2025; Zahit Demiray and Arslan, 2022). Prior research suggests that diagnostic value within a test sequence is often unevenly distributed, such that some steps contribute substantially more to defect detection than others, while other steps add limited or redundant value (Ferhani et al., 2008; Benner and Boroffice, 2001). To this end, Hamrol et al. (2020) proposed a value-based framework for evaluating quality inspection in multi-stage manufacturing processes. In their model, inspection effectiveness is assessed in terms of added value, defined as the difference between quality costs with and without inspection. Their results indicate that inspection at a given stage should be abandoned or improved if it does not generate sufficient value. This reinforces the broader argument that not all test or inspection steps contribute equally to manufacturing quality outcomes. In addition, Liang and Zheng (Liang and Zheng, 2022) present an industrial case study in semiconductor testing showing that optimizing the data extraction method within the test program can significantly reduce test time without altering the functional intent of the test sequence. Also, Iaria et al. (Iaria et al., 2025) demonstrate this clearly in large-scale automotive SoC production. Using a weighted fault coverage metric applied to 80 million faults across a 40 nm device, they show that the majority of defects are detected by a small fraction of available test patterns, whereas the remaining patterns contribute no additional diagnostic value, even after accounting for nonuniform defect distributions across the die.
In the specific context of PCBA manufacturing, even a small defect in a single component can affect the functionality of the entire board (Petkov and Ivanova, 2024). To detect such defects, PCBA production employs multiple sequential test stages (Li et al., 2023). In addition to traditional inspection methods, recent research has explored deep learning–based optical approaches to detect manufacturing defects and hidden hardware trojans in PCBAs, improving detection accuracy and coverage (Kulkarni and Xu, 2021). In this process, the PCBA stage includes mandatory firmware programming steps that must be executed on every unit, like electrical characterization steps that verify power supply voltages and operating currents, and functional verification steps that confirm communication bus connectivity and system-level board identity (Serban et al., 2014; Kis et al., 2019). The End-of- Line stage subsequently validates complete wireless system performance before shipment. Figure 1 illustrates the overall test flow evaluated in this work.
2.2 Data-driven test picking models
Several studies have proposed data-driven approaches to reduce test time by identifying which test steps provide meaningful diagnostic information and which may be safely omitted by (Ahsan et al., 2020; Karim Kausik et al., 2025). These approaches vary in how test reduction is applied, with some identifying a fixed reduced subset is identified offline and applied uniformly to all units, others using correlation rules used to customize the test sequence based on process measurements, and in some using adaptive policies that evolve over time as production results change (Rodrigues et al., 2013; Huang et al., 2016). Notably, many of these approaches rely on historical production data to develop the initial test-reduction logic before deployment, followed by limited adaptation during production (Song et al., 2022).
Specifically, for high-yield integrated circuit products, Pan et al. (Pan et al., 2025) proposed an ensemble-learning-based adaptive testing framework that addresses the severe class imbalance typical of such settings, where defective units represent a small fraction of production. In their framework, data re-balancing, feature selection, and decision boundary adjustment are combined to minimize the number of executed tests while maintaining high classification performance and limiting test escapes. Furthermore, Saha et al. (2025) proposed an adaptive testing framework for post-manufacturing testing of compute-in-memory CNN accelerators, in which test images are applied progressively and testing stops once the device can be classified as pass or fail with sufficient statistical confidence. A sequential estimation variant further improves efficiency by applying test images in an ordered sequence, achieving up to 4.6× speedup.
These models emphasize the need for adaptive test plan selection to handle shifts in production over time. Indeed, the selection of adaptive test plans has also been explored at the production level (Arslan and Orailoglu, 2011). For example, Rodrigues et al. (Rodrigues et al., 2013) proposed a multi-agent system for adaptive functional test plan customization in a real washing machine production line, in which quality and process data collected across production stages are correlated using an MPFQ model to adjust the sequence of functional tests for each appliance dynamically, thereby reducing functional test time through the elimination of unnecessary test steps. However, the adaptation mechanism relies on predefined correlation rules rather than learning from observed test outcomes, and it does not provide explicit risk-bounded guarantees on defect escape or cost-constrained test step reduction on a per-unit level. In general, these approaches remain connected to historical data or predefined decision rules, with limited capacity for continuous online learning from evolving production conditions.
2.3 Multi-armed bandit algorithms
MAB algorithms provide an intuitive framework for repeated decision-making under uncertainty (Slivkins, 2019; Vermorel and Mohri, 2005; Mahajan and Teneketzis, 2008). The classical analogy is that of a gambler facing several slot machines, each with an unknown payoff distribution, who must decide which machine to play over time. In each round, the decision-maker must balance two competing objectives: exploiting the option that currently appears most promising and exploring alternatives that may prove superior as more evidence is accumulated. This exploration–exploitation trade-off makes the MAB framework particularly suitable for sequential industrial decisions in which actions must be selected repeatedly while their true value is only gradually revealed through observed outcomes (Lazebnik, 2023; Lazebnik et al., 2024; Lazebnik et al., 2025; Russo et al., 2018; Bouneffouf et al., 2020). Figure 2 provides a schematic view of the MAB algorithms and their production tests prediction.
Formally, a stochastic MAB problem is defined over a finite set of arms A = {1, . . . , K}, where each arm represents a candidate decision and is associated with an unknown reward distribution ]a with mean μa = E[r | a] (Russo et al., 2018). At each time step t = 1, . . . , T, the agent selects an arm at ∈ A based on the accumulated history Ht−1 = {(as, rs)}t−1 s=1, observes a reward rt ~ ]at, and updates its internal estimate of arm quality. The goal is to learn a policy π that maximizes expected cumulative reward E[T t=1rt], or equivalently minimizes cumulative regret relative to the best arm in hindsight
where μ* = maxa∈Aμa. In the context of adaptive manufacturing testing, the arms may correspond to alternative test plans, while the reward function can be formulated to reflect the operational objective, for example, by rewarding test-time reduction and penalizing missed defect detections. Under this formulation, the exploration–exploitation trade-off becomes the problem of deciding when to apply a lower-cost reduced test plan and when to revert to a higher-cost but safer full test plan (Russo et al., 2018; Bouneffouf et al., 2020).
Several studies have demonstrated the usefulness of bandit methods in decision problems related to testing and quality control. In advanced manufacturing, Liu et al. (2023) proposed a context-aware combinatorial bandit framework for quality testing in a B5G-enabled production environment. Their method uses contextual production information to decide which products should be tested under limited testing capacity, showing that bandit learning can support adaptive quality-test allocation in non-stationary industrial settings. In software testing, Lima and Vergilio (2022) introduced the COLEMAN approach, which uses a MAB to prioritize test cases in continuous integration environments based on historical failure information. Their results show that bandit-based policies can adapt to volatile test environments in which test cases are added or removed over time. These studies are closely related to the present work because they demonstrate the suitability of bandit methods for sequential testing decisions under uncertainty. However, the present study considers per-unit selection between alternative manufacturing test plans while explicitly managing defect escape risk.
3 Task and model definition
We define the task over a set of available test steps, each with an associated execution cost. The goal is to identify a subset of these steps that preserves defect detection coverage while reducing the total cost of testing each unit (Iaria et al., 2025). Formally, let T = {1, . . . , M} denote the index set of available tests in a given stage, and let ci > 0 be the execution cost (e.g., mean time) of test i ∈ T . The total cost of executing all test steps is cfull = i∈T ci. This represents the per-unit test time under the current static test plan, which serves as the baseline cost against which all reductions are measured.
To capture the diagnostic behaviour of each step, consider N historical units tested with the full flow. Let us define a binary outcome matrix Y ∈ {0, 1}N×M with entries:
Each row of Y represents the complete test outcome profile of one historical unit, and each column represents the pass/fail history of one test step across all units. Moreover, let UF = { u: ∃i ∈ T s.t. yu,i = 0 } be the set of historically failing units. In plain terms, UF contains every unit that failed at least one test step during the historical observation period, excluding equipment-induced failures. For any candidate subset C ⊆ T , a failing unit u ∈ UF is detected by C if at least one executed test in C fails historically DC(u) = maxi∈C(1 −yu,i) ∈ {0, 1}. Simply put, DC(u) = 1 means subset C would have caught defective unit u; DC(u) = 0 means unit u would have escaped detection under subset C. This definition also handles linked or cascading failures, in which a single underlying product defect causes multiple test steps to fail on the same unit. The coverage objective is defined at the unit level rather than at the failed-test level: a defective unit is counted once, and it is considered detected if at least one of its failed test steps is included in C. Therefore, multiple simultaneous failures on the same unit do not artificially increase the number of covered defects. Instead, they appear as redundant diagnostic evidence for the same unit.
| Symbol | Definition |
|---|---|
| T = {1, . . . , M} | Index set of all available test steps |
| M | Total number of test steps |
| ci | Execution cost (mean time in seconds) of test step i |
| cfull | Total cost of executing all test steps |
| N | Number of historical units in training dataset |
| Y ∈{0, 1}N×M | Binary outcome matrix |
| yu,i | Outcome of unit u on test step i (1 = pass, 0 = fail) |
| UF | Set of historically failing units |
| C ⊆T | A candidate subset of test steps |
| Cfull | The complete set of all test steps T |
| Cred | The reduced subset |
| DC(u) | Detection indicator for unit u under subset C |
| ˆR(C) | Empirical escape risk of subset C |
| ϵ | Tolerable escape risk threshold |
| δ | Maximum allowed escape risk for MAB policy |
| π | Test plan selection policy |
| ρt | Rolling pass rate at unit t |
| w | Rolling window size |
| τ | Process stability threshold |
| β | Instability sensitivity parameter |
| κ | Escape penalty in the reward function |
| αa, βa | Beta distribution parameters for arm a |
| θa | Sampled value for arm a |
| rt | Reward signal at unit t |
| at | Arm selected at unit t |
Furthermore, the empirical escape risk of C is the fraction of historically failing units not detected by C, defined as ˆR(C) = |UF|u∈UF1{DC(u) = 0}. A value of 1 ˆR(C) = 0 means subset C detects all historically failing units with zero escapes, while a value of ˆR(C) = 1 means all historically failing units would escape detection under subset C.
Based on this formalization, we defined three tasks: test value assessment, risk-bounded reduction, and dynamic test strategy. First, find the cheapest set of test steps that catches every defective unit observed in the training data:
Second, find the minimum-cost subset under a tolerable empirical escape-risk threshold ϵ ∈ [0, 1]:
Third, let Cfull = T and Cred ⊆ T . A policy π: [0, 1] → {Cfull, Cred} maps an observed process stability signal ρt ∈ [0, 1] to a test plan at each unit t. The goal is to minimize expected cost while bounding the policy-induced empirical escape risk by δ ∈ [0, 1]:

where ˆR(π) = |UF| u∈UF1{Dπ(ρt)(u) = 0} is computed by replaying 1 policy decisions over historical outcomes. To this end, unlike the first two tasks, which produce static subsets from historical data, this task requires a policy that adapts in real time to non-stationary production conditions, a sequential decision problem under uncertainty (Bouneffouf et al., 2020; Lazebnik, 2026).
Table 1 summarizes the notations with their definitions.
With this formulation established, the proposed framework addresses the three research questions through two sequential phases. The offline phase solves the first two tasks using historical production data to construct the reduced test subset Cred and characterize the cost-risk trade-off. In a complementary manner, the online phase solves the third task by deploying a Thompson Sampling MAB-based model that dynamically selects between Cfull and Cred for each incoming unit based on a real-time process stability signal (Russo et al., 2018; Nie, 2024; Han et al., 2025).
Formally, for the first task, we use the historical outcome matrix Y and cost vector c prior to deployment. These two stages construct the reduced subset Cred that serves as the fixed alternative arm in the MAB agent. At each iteration, the greedy algorithm selects the test step that provides the greatest additional defect detection per unit of execution cost:
where {u ∈ UF: yu,i = 0} is the set of failing units that step i would detect, Covered(C) is the set of failing units already detected by steps already in C, and the numerator counts only the newly detected failing units that step i would add. This newly-covered-unit criterion is important when failures are linked across tests. If several tests fail because of the same underlying defect, then after one of those tests has been selected, the remaining linked tests contribute little or no additional coverage unless they also detect other failing units. Hence, the greedy procedure naturally avoids selecting multiple tests that are redundant with respect to the same linked failure pattern. Dividing by ci ensure cheaper steps are preferred when they provide equivalent detection. Importantly, since the weighted set cover problem is NP-hard (Adamo et al., 2023), an exact solution is intractable for large test suites. The greedy algorithm provides an approximation guarantee of 1 −1/e ≈ 63% of the optimal solution in polynomial time, making it practical for industrial test suites with hundreds of steps (Chvátal, 1979; Adamo et al., 2023; Iaria et al., 2025). The output is the minimum-cost subset Cred = C* that satisfies ˆR(C*) = 0.
Next, the second task extends the first task by sweeping the allowed escape threshold ϵ from 0 to |UF|, applying the greedy cover at each level to generate the full Pareto frontier of cost savings versus escape risk. Namely, the frontier F is initialized as empty. For each escape level ε ∈ {0, 1, . . . , |UF|}, Algorithm 1 is applied with the escape tolerance set to ε/|UF|, producing a subset Cε. The cost saving and empirical escape risk of Cε are then computed. Afterwards, a Pareto dominance check is performed inwhich the point is added to F only if no existing point in F simultaneously achieves both higher saving and lower escape risk. This ensures that F contains only non-dominated operating points, each representing the minimum-cost subset achievable at its corresponding escape tolerance level.
Finally, the last task is cast as a two-armed stochastic MAB. The two arms correspond to the full test plan Cfull (arm 0) and the reduced test plan Cred (arm 1). At each production unit t, the agent selects one arm, executes the corresponding test plan, observes the outcome, and updates its belief about the value of each arm (Bouneffouf et al., 2020). Before each arm selection decision, the agent computes a rolling pass rate over the last w units:
where w is the window size and 1[outcome(ui) = PASS] equals 1 only if unit ui passed all executed steps, and 0 for either explicit product failures or abort-only executions. Thus, abort-only events are not used as genuine product defects when constructing UF, but they still reduce the rolling pass-rate signal ρt during online operation. Consequently, a temporary increase in equipment-induced aborts is interpreted by the policy as process instability and biases the Thompson Sampling decision toward Cfull. This behavior is conservative: it can reduce the realized test-time saving, but it does not encourage additional use of Cred under unstable conditions. The rolling pass rate ρt serves as a proxy for process stability: a high value (close to 1) indicates stable production where the reduced plan is likely safe; a low value signals elevated defect risk where the full plan should be preferred. The process is classified as stable if ρt ≥τ, where τ is a pre-specified stability threshold. In the experiments, we set τ = 0.93 as the nominal stability threshold. This value was chosen to impose a conservative definition of process stability: reduced testing is permitted only when the recent rolling pass rate remains close to one, while relatively small degradations in pass rate trigger a shift toward the full test plan. Operationally, τ therefore controls the quality-cost trade-off of the online policy. Lower values of τ make the policy more permissive and increase the use of Cred, whereas higher values make the policy more conservative and increase the use of Cfull. When the process is unstable (ρt < τ), the Thompson Sampling score for Cfull is boosted by an instability penalty ˜θ0 = θ0 + β · (τ −ρt), where θ0 is the sampled value for arm 0, τ −ρt > 0 is the magnitude of the instability, and β ≥0 is the instability sensitivity parameter. A larger β makes the agent more conservative, reverting to Cfull more readily when instability is detected. When the process is stable (ρt ≥τ), no boost is applied and ˜θ0 = θ0. Of note, this reward signal is designed to align agent behaviour with the quality-cost objective (Nie, 2024; Liu et al., 2023). The reward values reflect a strict quality hierarchy where defect escapes are penalized most severely, cost savings on clean units are rewarded maximally, and defect detection under the reduced plan receives a partial reward to acknowledge maintained quality coverage at reduced cost. For a unit ut processed under arm at ∈ {0, 1}, the reward is defined as:

where κ > 0 is the escape penalty. The reward values are normalized utilities chosen to encode the priority order of the decision problem rather than direct monetary costs. Selecting Cfull receives a neutral reward because it preserves quality but provides no time saving. Selecting Cred on a defect-free unit receives the maximum reward because it realizes the full test-time saving. Selecting Cred when a defect is still detected receives an intermediate reward because quality is preserved, but the observation indicates a higher-risk production state. An escaped defect receives the negative reward −κ, making escape events dominate the positive utility accumulated from successful reduced-test executions. Selecting Cfull yields a neutral reward of zero–it is safe but costly. Selecting Cred on a clean unit yields the maximum reward. A detected defect under Cred yields a partial reward, signalling that quality was maintained but risk was present. An escape yields a large negative reward, discouraging the agent from selecting Cred under high-risk conditions. importantly, κ was treated as a conservatism parameter of the online policy. It was selected a priori to be larger than the maximum positive reward, so that a single escape event outweighs multiple successful uses of the reduced test plan. As a result, increasing κ makes the Thompson Sampling agent more conservative since after an escape, the posterior value of Cred is reduced more strongly and the policy shifts more readily toward Cfull. Thus, decreasing κ makes the policy more aggressive, increasing the expected use of Cred and the potential saving, but also increasing exposure to defect escapes under drift.
Figure 3 presents the schematic flow of the complete framework, from offline subset construction to online adaptive test execution, illustrating how the two phases interact to continuously adapt test plan selection to changing production conditions without requiring offline retraining (Russo et al., 2018; Bouneffouf et al., 2020).
4 Experiments
To evaluate the proposed model, we first outline a real-world experiment design and then present the obtained results of our analysis.
4.1 Experiment design
The experiments are conducted in collaboration with an industrial partner operating a high-volume electronics manufacturing line, where Printed Circuit Board Assemblies are produced and tested as part of the manufacturing process. In the current production process, every board undergoes two mandatory sequential test stages before integration into the final product: a Functional Circuit Test stage comprising 172 individual test steps that verify board-level electrical functionality and firmware integrity, and an End-of-Line stage comprising 55 test steps that validate complete wireless system performance. Any board failing either stage is rejected from the production flow. Under the current static test plan, the full sequence of steps is executed for every produced unit regardless of its production history or the current state of the production process. The mean full-sequence test time is 157.62 s per unit at FCT and 88.00 s per unit at EOL, representing a substantial proportion of the total per-unit production time. The production line operates under a strict quality threshold of δ = 8.5 × 10−5 defect escapes per unit, which must be maintained at all times. However, executing the full test sequence for every unit regardless of current production conditions results in significant test time that can be reduced without compromising defect detection coverage under stable production conditions.
Formally, in this study, we use two datasets covering two structurally different test stages. Both datasets consist of raw test log files exported from the LabVIEW-based production test equipment, with one structured text file per board run recording the pass or fail outcome of each test step along with its execution time. The first dataset covers the Functional Circuit Test stage, comprising 172 distinct test steps executed across 7,618 board runs during the training period and 6,081 board runs during the validation period. Of the 332 board-level failures in the FCT training set, 302 are genuine defect events and 30 are abort-only runs attributed to test equipment errors rather than product defects, corresponding to a genuine defect rate of 3.96% across the training period. The second dataset covers the End-of-Line test stage, comprising 55 distinct test steps executed across 10,633 board runs during training and 4,585 board runs during validation. Of the 416 board-level failures in the EOL training set, 318 are genuine defect events and 98 are abort-only runs, corresponding to a genuine defect rate of 2.99%. The mean full-sequence test time is 157.62 s per board for FCT and 88.00 s per board for EOL.
Raw test log files were exported as structured text files, with one file per board run, and merged into a single flat CSV (Comma Separated Values) per dataset (van den Burg et al., 2019). Column names were normalized to lowercase with underscores, and header-repeat rows-artifacts of the log export format were removed as part of standard data normalization and cleaning (Sankpal and Metre, 2020). Each unique source filename was treated as a distinct unit identifier, as each file corresponds to one independent board run. A unique test step identifier was constructed by concatenating two hierarchical name fields, producing 172 unique step identifiers for FCT and 55 for EOL. Step execution times were cast to numeric values with missing entries filled as zero. Moreover, units whose board-level outcome was FAIL but whose individual step results contained no explicit FAIL, only ABORTs were classified as abort-only and excluded from the genuine failing unit set UF. This filtering approach follows established practice in production test data cleaning, where equipment-induced failures are distinguished from genuine product defects prior to analysis. This exclusion removed 30 units from FCT and 98 units from EOL. An abort-only unit is therefore not treated as evidence of a diagnosable product defect, because no individual test step produced an explicit product-related FAIL. At the same time, abort-only records may contain incomplete diagnostic information and could, in principle, mask a latent product defect if the test sequence terminated before reaching the relevant diagnostic step. For this reason, these units were excluded only from the genuine defect set UF used for subset construction and empirical escape-risk calculation; they should not be interpreted as passing units. In an operational deployment, abort-only units would be routed to retest or engineering review rather than released as conforming products. The remaining 302 FCT and 318 EOL genuine failing units formed UF for all subsequent computations. In addition, a cost vector c ∈ RM of mean step execution times was derived by averaging step execution times across all runs for each step (Liang and Zheng, 2022; Zahit Demiray and Arslan, 2022). The FCT data set comprises 7,618 training runs and 6,081 validation runs across 172 test steps, with 302 genuine defect events in training and 868 in validation a 3.6 × increase in defect rate. The EOL dataset comprises 10,633 training runs and 4,585 validation runs across 55 test steps, with 318 genuine defect events in training and 93 in validation. Both datasets are partitioned using a temporal split such that July-December 2025 forms the training set and January- February 2026 forms the held-out validation set.
In order to evaluate the proposed model’s performance, four metrics are used and reported across both training and validation periods. Together they capture the two competing objectives of the framework: reducing test time and maintaining defect detection coverage. Table 2 summarizes these four evaluation metrics in terms of notation, mathematical formalization, and motivation.
The proposed model is evaluated in three sequential stages. First, a greedy set cover algorithm (Iaria et al., 2025) is applied to the training outcome matrix to identify the minimum-cost subset Cred ⊆ T that detects all genuine training failures with zero escape risk. At each iteration, the step with the highest ratio of newly detected failing units to execution cost is selected until all failing units are covered. Next, the Pareto frontier of cost saving versus escape risk is constructed by sweeping the allowed escape threshold from zero to |UF|, generating one Pareto-optimal subset per unique cost-saving level (Pan et al., 2025). For the FCT dataset, this is done using a dense sweep over all integer escape thresholds. For the EOL dataset, all 214 possible subsets of diagnostic steps are enumerated exhaustively because the number of steps is small enough to allow such analysis. Finally, three MAB policies are trained online over the chronological unit stream: Thompson Sampling (Russo et al., 2018), Upper Confidence Bound (UCB) (Bouneffouf et al., 2020), and Epsilon-Greedy (Nie, 2024). Each policy is initialized with a conservative prior favoring Cfull (α0 = 5, β0 = 1) and a uniform prior for Cred (α1 = 1, β1 = 1). The escape penalty κ was fixed across all experiments so that the reported differences reflect the adaptive policy response to process instability rather than retuning of the reward function. The rolling pass-rate stability signal is computed over a sliding window of w consecutive units and compared against threshold τ before each arm-selection decision. Thompson Sampling is then evaluated on the held-out validation stream under multiple instability sensitivity values (β ∈ {10, 50} for FCT and β ∈ {10, 50, 100} for EOL) to characterize the quality-cost trade-off under concept drift. In addition, we perform a sensitivity analysis of the stability threshold by repeating the validation experiment with τ ∈ {0.90, 0.93, 0.95}. This analysis evaluates whether the conclusions depend on the nominal threshold τ = 0.93 and quantifies the effect of a more permissive or more conservative stability criterion.
| Test | Cohort | Algorithm | β | Cred (%) | Saving (%) | Escaped | ˆR(π) |
|---|---|---|---|---|---|---|---|
| FCT | Training | Baseline (Cfull) | — | 0.0 | 0.00 | 0 | 0.00 |
| Static Cred | — | 100.0 | 18.78 | 0 | 0.00 | ||
| Epsilon-greedy | 10 | 89.6 | 16.83 | 0 | 0.00 | ||
| UCB | 10 | 92.8 | 17.42 | 0 | 0.00 | ||
| Thompson sampling | 10 | 93.8 | 17.61 | 0 | 0.00 | ||
| Validation | Baseline (Cfull) | — | 0.0 | 0.00 | 0 | 0.00 | |
| Static Cred | — | 100.0 | 20.45 | 110 | 0.13 | ||
| Thompson sampling | 10 | 16.1 | 3.30 | 8 | 0.01 | ||
| Thompson sampling | 50 | 2.4 | 0.49 | 0 | 0.00 | ||
| EOL | Training | Baseline (Cfull) | — | 0.0 | 0.00 | 0 | 0.00 |
| Static Cred | — | 100.0 | 91.57 | 0 | 0.00 | ||
| Epsilon-greedy | 10 | 87.0 | 79.68 | 0 | 0.00 | ||
| UCB | 10 | 90.9 | 83.24 | 0 | 0.00 | ||
| Thompson sampling | 10 | 91.5 | 83.77 | 0 | 0.00 | ||
| Validation | Baseline (Cfull) | — | 0.0 | 0.00 | 0 | 0.00 | |
| Static Cred | — | 100.0 | 91.57 | 8 | 0.09 | ||
| Thompson sampling | 10 | 98.2 | 89.95 | 1 | 0.01 | ||
| Thompson sampling | 50 | 97.6 | 89.35 | 1 | 0.01 | ||
| Thompson sampling | 100 | 95.2 | 87.17 | 0 | 0.00 |
| Test stage | τ | Cred (%) | Saving (%) | Escaped | ˆR(π) |
|---|---|---|---|---|---|
| FCT | 0.90 | 14.8 | 3.03 | 6 | 0.007 |
| FCT | 0.93 | 2.4 | 0.49 | 0 | 0.000 |
| FCT | 0.95 | 0.9 | 0.18 | 0 | 0.000 |
| EOL | 0.90 | 98.8 | 90.47 | 2 | 0.022 |
| EOL | 0.93 | 95.2 | 87.17 | 0 | 0.000 |
| EOL | 0.95 | 88.0 | 80.58 | 0 | 0.000 |
4.2 Results
Table 3 summarizes the results for both the FCT and EOL datasets across training and validation cohorts. In both datasets, all MAB methods achieve zero escapes during training while approaching the savings of the static reduced plan. Under validation, however, the static Cred policy fails due to concept drift, producing 110 escapes on FCT and 8 on EOL. In contrast, Thompson Sampling adapts by shifting back toward the full plan when instability is detected, reducing escapes to zero at β = 50 for FCT and β = 100 for EOL. This shows that dynamic test selection preserves most of the cost benefit of reduction during stable periods while remaining robust to unseen failure modes at deployment.
Table 4 reports the sensitivity of the adaptive policy to the stability threshold. As expected, decreasing τ to 0.90 makes the policy more permissive, increasing the fraction of units assigned to Cred and therefore increasing test-time savings. However, because reduced testing is allowed during weaker stability conditions, this setting also increases the probability of defect escapes under concept drift. Increasing τ to 0.95 has the opposite effect: the policy reverts to Cfull more frequently, reducing escape risk but lowering the achievable saving. The nominal value τ = 0.93 therefore represents an intermediate operating point that preserves the main quality objective while avoiding the excessive conservatism of a higher threshold.
Figure 4 shows the cost-risk Pareto fronts for both test suites. Panel (a) presents the FCT frontier, where the zero-escape operating point achieves 18.78% saving using a 32-step subset that removes 140 redundant steps. Across all 13,699 FCT board runs, every step-level failure also produced a board-level failure, indicating that the rolling pass-rate signal captures all observed failure modes, including those outside Cred. The Pareto frontier follows a clear logarithmic trend, Ŝ(ε) = 15.30 ln(ε) + 106.34 with R2 = 0.994, showing strongly diminishing returns: most achievable savings are concentrated in the low-risk region near the operating point. Because coverage is computed per defective unit, these savings are not obtained by counting multiple linked test failures as separate defects. A unit with several simultaneous failed steps contributes only one coverage requirement, and the reduced subset must include at least one test capable of detecting that unit. Panel (b) presents the EOL frontier, obtained by exhaustively evaluating all 214 subsets of diagnostic steps. Unlike FCT, the frontier exhibits visible discrete banding because only 14 diagnostic steps are available, so achievable saving levels are quantized. Despite this, the zero-escape operating point reaches 91.57% saving, substantially higher than in FCT, because the 41 non-diagnostic EOL steps account for most of the total test time.
Figure 5 shows the rolling fraction of training units assigned to Cred by the Thompson Sampling agent for both datasets. Panel (a) presents the FCT training stream over 7,618 units. The agent converges to about 95% Cred selection within the first 500 units, then drops sharply to nearly 35% during the detected instability region before recovering, showing that the stability signal can temporarily override the learned cost preference when quality risk increases and that this response is reversible. Panel (b) presents the EOL training stream, where the agent converges to approximately 95% Cred selection by unit 1,500. Unlike FCT, which contains a single pronounced instability event, EOL exhibits multiple scattered instability regions across the training period, indicating a less stable production process. The agent responds to each episode by reducing Cred selection and recovering afterward, demonstrating robust adaptive behavior under varying process conditions.
Figure 6 summarizes the rolling pass rate and Cred selection behavior during validation for both datasets. Panels (a) and (b) show the FCT validation stream. In contrast to the training period, which contained only one isolated instability event, the validation period is persistently unstable, with the rolling pass rate fluctuating between 0.60 and 0.92 and remaining below the stability threshold τ = 0.93. As a result, the agent automatically suppresses Cred selection to below 30% without retraining, reflecting a conservative response driven entirely by the stability signal. Panels (c) and (d) show the EOL validation stream. Here, the rolling pass rate fluctuates around the same threshold rather than remaining consistently below it, producing a more dynamic behavior: the agent retains high overall Cred selection while still reacting to individual instability events. This enables zero escapes at β = 100 while preserving 87.17% cost saving.
5 Discussion
In this study, we proposed an MAB-based model for quality-preserving production test reduction. We evaluate the proposed model on a 2-stage test set from a real-world production line of electronics. The results demonstrate that the value of the proposed framework lies not only in reducing test effort, but in doing so while preserving quality under changing production conditions. This is most clearly seen in the validation results reported in Table 3. In the FCT stage, the static reduced plan remained highly effective on the training period, achieving 18.78% test-time saving with zero escapes, but failed under temporal validation with concept drift, producing 110 escaped defects. By contrast, the adaptive Thompson Sampling policy reduced this to 8 escapes at β = 10 and to zero escapes at β = 50, albeit with a corresponding reduction in cost saving. A similar pattern is observed in the EOL stage, where the static reduced plan achieved 91.57% saving but produced 8 escapes in validation, whereas the adaptive policy achieved zero escapes at β = 100 while still preserving 87.17% saving. The dynamics underlying these aggregate outcomes are visible in Figure 6, which shows that the policy suppresses selection of Cred when the rolling pass-rate signal indicates instability and resumes reduced testing when conditions recover. The threshold-sensitivity analysis in Table 4 further shows that this behavior is not tied to a single arbitrary value of τ: lower thresholds favor savings but increase exposure to drift, whereas higher thresholds favor quality protection at the cost of reduced savings.
These findings are consistent with the broader shift in manufacturing quality engineering from static inspection design toward closed-loop, data-driven quality management. Recent work on zero-defect manufacturing argues that modern quality systems should combine defect detection, defect prevention, and adaptive decision support across the production chain rather than rely solely on fixed downstream inspection policies (Powell et al., 2022). Likewise, digital quality-management platforms have been proposed to integrate production, process, and quality data in ways that support continuously updated operational decisions rather than isolated offline analyses (Filz et al., 2024). Viewed in that context, the present framework is not merely a local optimization of testing effort. Rather, it can be interpreted as a practical mechanism for translating unit-level production evidence into adaptive quality-control actions during operation.
The Pareto fronts in Figure 4 make the trade-off of appraisal effort and failure-related cost explicit. In FCT, the frontier shows that a substantial proportion of the achievable saving is concentrated near the zero-escape operating point, indicating that moderate time reduction is possible before defect risk rises sharply. In EOL, the frontier is much steeper and more discrete, reflecting the smaller number of diagnostically relevant steps and the much larger share of non-diagnostic test time. This interpretation is aligned with prior work on multistage inspection planning, which has shown that inspection efficiency depends on the combined effect of process capability, inspection cost, and downstream non-conformance cost (Hamrol et al., 2020). It is also consistent with cost-of-quality analyses demonstrating that the economically preferable inspection strategy is generally the one that balances appraisal savings against the risk of repair, scrap, and escaped failures, rather than the one that simply minimizes inspection time (Arsalan Farooq et al., 2017).
The training-stream behavior shown in Figure 5 further clarifies the role of the adaptive agent. In both stages, the policy learns to favor the reduced plan during stable operating regions, but it does not do so monotonically or irreversibly. Instead, it falls back toward the full plan when the process-stability signal deteriorates, then returns to higher reduced-plan usage when stability is restored. This behavior is important from an industrial standpoint because it suggests that reverting to Cfull should not be interpreted as a weakness of the method, but as the intended safety response of a learning policy deployed in a quality-critical environment. In this sense, the framework is closely related to conservative online-learning formulations, in which exploration is permitted only while maintaining performance relative to a trusted baseline policy (Wu et al., 2016; Kazerouni et al., 2017). In the present application, the full test plan serves as that trusted baseline.
From an application perspective, the results suggest that deployment should proceed gradually and within existing manufacturing quality infrastructures. A practical first step would be shadow-mode deployment, in which the learned policy recommends a plan while the executed plan remains under engineering control. Such a deployment strategy is consistent with recent work on data-driven quality platforms and real-time hybrid inspection systems, both of which emphasize traceability, staged integration, and the use of predictive models to complement rather than abruptly replace established inspection procedures (Filz et al., 2024; Mohamed et al., 2022). In addition, the reward structure of the MAB agent should ultimately be calibrated in plant-specific economic terms, so that the relative utility assigned to test-time reduction, defect detection, rework, scrap, and escape events reflects the actual cost-of-quality structure of the production environment. This is especially important for the escape penalty κ, which controls the conservatism of the learned policy. A larger κ is appropriate in safety-critical or warranty-sensitive production settings where escaped defects are extremely costly, whereas a smaller κ may be acceptable in lower-risk settings where additional test-time reduction is prioritized.
This study is not without limitations. First, the framework is evaluated on a single industrial PCBA setting, and broader validation across additional products, production lines, manufacturing technologies, and defect regimes will be required to establish the generality of the observed behavior. The results therefore demonstrate feasibility and practical value in the studied factory, but they should not yet be interpreted as universal performance guarantees for all manufacturing environments. Second, the online decision problem is formulated as a binary choice between Cfull and a single reduced subset, whereas many production settings may benefit from richer action spaces involving multiple Pareto-optimal subsets or stage-specific test intensities. Third, the reward function compresses multiple operational objectives into a scalar signal, which is convenient for online learning but may under-represent rare, high-consequence escapes. Fourth, the current adaptation mechanism relies on the rolling pass-rate signal as its primary indicator of changing production conditions; although this proved effective in the studied datasets, richer contextual signals from upstream process measurements, material batches, intermediate quality states, or equipment-health indicators may improve responsiveness under subtle forms of drift. In particular, abort-only runs caused by tester or fixture interruptions can reduce the rolling pass rate in the same way as product-quality failures. The resulting response is conservative because the policy shifts toward Cfull, but it may conflate equipment instability with product-quality drift and thereby reduce achievable savings. A practical deployment should therefore monitor abort-only rates separately and use equipment-status features to distinguish tester instability from genuine process drift. Fifth, the current formulation assumes fixed average test costs. In practice, test duration may change over time because of machine slowdowns, fixture degradation, calibration state, queueing delays, operator interventions, or other equipment-health factors. Future work should therefore extend the cost model from fixed mean costs ci to dynamic costs ci(t) or context-dependent costs ci(xt), allowing the policy to account for changing equipment conditions during production. In addition, although the set-cover formulation uses the observed pass/fail matrix and therefore captures empirical co-failure patterns, it does not explicitly model causal dependencies among tests. When one failure mode causes several tests to fail together, the current method treats this as a unit-level detection pattern rather than estimating a root-cause or conditional-dependence structure. Future extensions could combine the proposed framework with fault-tree models, Bayesian networks, or causal diagnostic models to represent dependencies among tests more explicitly. Finally, industrial adoption depends not only on performance but also on interpretability, and future work should therefore consider explanation mechanisms that clarify why the policy selected the reduced or full plan for a given unit.
Taken jointly, the results show that adaptive test selection offers a practical middle ground between fully conservative inspection and static test reduction. The principal contribution of the study is therefore not simply that it reduces test time, but that it demonstrates how online learning can be aligned with manufacturing quality objectives in a way that remains operationally cautious under concept drift. By combining offline subset optimization with online policy adaptation, the framework provides a credible path toward more efficient manufacturing test operations without abandoning the central requirement of controlling defect escape.
Data availability statement
The datasets analyzed in this study are not publicly available because they contain proprietary industrial production test data. Requests to access the datasets should be directed to the corresponding author and will be considered subject to confidentiality restrictions and approval by the industrial partner.
Author contributions
EP-A: Conceptualization, Supervision, Writing – original draft, Writing – review and editing. NH: Conceptualization, Data curation, Formal Analysis, Methodology, Software, Visualization, Writing – original draft, Writing – review and editing. TL: Conceptualization, Formal Analysis, Investigation, Supervision, Validation, Visualization, Writing – original draft, Writing – review and editing.
Funding
The author(s) declared that financial support was received for this work and/or its publication. This research was funded by Vinnova under grant number 2026-00166 within the project “BELIEF -From prediction to trust: AI-based decision support for maintenance”.
Conflict of interest
The author(s) declared that this work was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.
Generative AI statement
The author(s) declared that generative AI was used in the creation of this manuscript. The authors used AI for code analysis generation, code documentation, initial related work search, schematic figures preparation, and the initial version of the manuscript. The authors declare that the final content is manually audited, and the authors take full responsibility.
Any alternative text (alt text) provided alongside figures in this article has been generated by Frontiers with the support of artificial intelligence and reasonable efforts have been made to ensure accuracy, including review by the authors wherever possible. If you identify any issues, please contact us.
Publisher’s note
All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.
Article notes
- Publication history
- Received 21 April 2026 · Accepted 25 May 2026 · Published 17 June 2026
- Keywords
- adaptive test selection
- multi-armed bandits
- Thompson sampling
- electronics manufacturing
- PCBA testing
- concept drift
References
- Adamo, T., Ghiani, G., Guerriero, E., and Pareo, D. (2023). A Surprisal-based Greedy Heuristic for the Set Covering Problem.
- Ahsan, M., Stoyanov, S., Bailey, C., and Albarbar, A. (2020). Developing computational intelligence for smart qualification testing of electronic products. IEEE Access 8, 16922–16933. Antonio do Prado Lima, J., and Vergilio, S. (2020). Multi-armed bandit test case prioritization in continuous integration environments: a trade-off analysis. 10 . doi:10.1109/access.2020.2967858
- Anusuya, M., Kavitha, P., Bathrinath, S., Vundrajavarapu, P., Bharath Kumar, R., and Sakthivel, M. (2024). “Automation of test and measurement in electronics manufacturing through ai,” in 2024 International Conference on Science Technology Engineering and Management (ICSTEM), 1–6.
- Arsalan Farooq, M., Kirchain, R., Novoa, H., and Araújo, A. (2017). Cost of quality: evaluating cost-quality trade-offs for inspection strategies of manufacturing processes. Int. J. Prod. Econ. 188, 156–166. Arslan, B., and Orailoglu, A. (2011). “Adaptive test framework for achieving target test quality at minimal cost,” in 2011 Asian Test Symposium, 323–328. doi:10.1016/j.ijpe.2017.03.019
- Benner, S., and Boroffice, O. (2001). “Optimal production test times through adaptive test programming,” in Proceedings International Test Conference 2001 (Cat. No.01CH37260), 908–915.
- Bhowmick, A., and Seetharaman, A. (2023). “Impact of product quality on customer satisfaction: a systematic literature review,” in Proceedings of the 2023 7th International Conference on Virtual and Augmented Reality Simulations, 93–99.
- Boumen, R., de Jong, I. S. M., Vermunt, J. W. H., van de Mortel-Fronczak, J. M., and Rooda, J. E. (2008). Test sequencing in complex manufacturing systems. IEEE Trans. Syst. Man, Cybern. - Part A Syst. Humans 38 (1), 25–37. Bouneffouf, D., Rish, I., and Aggarwal, C. C. (2020). “Survey on applications of multi-armed and contextual bandits,” in 2020 IEEE Congress on Evolutionary Computation (CEC), 1–8. doi:10.1109/tsmca.2007.909494
- Cecchitelli, M., Fiori, G., Galo, J., Andrea Sciuto, S., and Scorza, A. (2026).
- Measurements for quality control in biomedical 3-d printing: standards, gaps, and missing protocols. IEEE Sensors Rev. 3, 243–256. Cerquitelli, T., Pagliari, D. J., Calimera, A., Bottaccioli, L., Patti, E., Acquaviva, A., et al. (2021). Manufacturing as a data-driven practice: methodologies, technologies, and tools. Proc. IEEE 109 (4), 399–422. doi:10.1109/jproc.2021.3056006 Charan Kantumchu, V., Quadir Moinuddin, S., Kumar Dewangan, A., and Cheepu, M. (2024). Quality Assurance and Control in Welding and Additive Manufacturing, 245–261. doi:10.1109/sr.2026.3657720
- Chiara Magnanini, M., Demir, O., Colledani, M., and Tolio, T. (2024). Performance evaluation of multi-stage manufacturing systems operating under feedback and feedforward quality control loops. CIRP Ann. 73 (1), 349–352. 2024.04.015 Chowdhury, A., and Nuruzzaman, Md (2023). Design, testing, and troubleshooting of industrial equipment: a systematic review of integration techniques for us manufacturing plants. Rev. Appl. Sci. Technol. 2 (01), 53–84. doi:10.63125/893et038 Chvátal, V. (1979). A greedy heuristic for the set-covering problem. Math. Oper. Res. 4, 233–235. doi:10.1287/moor.4.3.233 Duarte, J. G., Duarte, M. G., Piedade, A. P., and Mascarenhas-Melo, F. (2025). doi:10.1016/j.cirp
- Rethinking pharmaceutical industry with quality by design: application in research, development, manufacturing, and quality assurance. AAPS Journal 27 (4), 96. doi:10. 1208/s12248-025-01079-w Fan, H., Liu, X., Fuh, J. Y. H., Wen, F.Lu, and Li, B. (2025). Embodied intelligence in manufacturing: leveraging large language models for autonomous industrial robotics. J. Intelligent Manuf. 36 (2), 1141–1157. Ferhani, F.-F., Saxena, N. R., McCluskey, E. J., and Nigh, P. (2008). “How many test patterns are useless?,” in 26th IEEE VLSI Test Symposium (Vts 2008), 23–28. doi:10.1007/s10845-023-02294-y
- Filz, M.-A., Bosse, J. P., and Herrmann, C. (2024). Digitalization platform for data-driven quality management in multi-stage manufacturing systems. J. Intelligent Manuf. 35 (6), 2699–2718. Ghelani, H. (2024a). Ai-driven quality control in pcb manufacturing: enhancing production efficiency and precision. Val. Int. J. Digital Libr. 12 (10), 1549–1564. doi:10.18535/ijsrm/v12i10.ec06 Ghelani, H. (2024b). Advanced ai technologies for defect prevention and yield optimization in pcb manufacturing. Int. J. Of Eng. And Comput. Sci. 13 (10), 26534–26550. doi:10.18535/ijecs/v13i10.4924 Grznár, P., Papánek, L., Marčan, M., Krajčovič, M., Antoniuk, I., Mozol, Š., et al. (2025). doi:10.1007/s10845-023-02162-9
- Enhancing production efficiency through digital twin simulation scheduling. Appl. Sci. 15 (7), 3637. Hamrol, A., Kujawińska, A., and Bożek, M. (2020). Quality inspection planning within a multistage manufacturing process based on the added value criterion. Int. J. Adv. Manuf. Technol. 108, 1–14. doi:10.1007/s00170-020-05453-0 Han, X., Cai, Y., Zhang, A., Zhu, Y., He, Y., and Shi, R. (2025). “Adaptive policy optimization for product infant failure risk control in manufacturing processes with variable demands,” in 2025 16th International Conference on Reliability, Maintainability and Safety (ICRMS), 376–381. doi:10.3390/app15073637
- Huang, Ke, Wen, J., and Willmore, J. (2016). Test-suite-based analog/rf test time reduction using canonical correlation. IEEE Trans. Computer-Aided Des. Integr. Circuits Syst. 35 (12), 2143–2147. Iaria, G., Bernardi, P., Bertani, C., Cardone, L., Garozzo, G., and Tancorre, V. (2025). A comprehensive scan test cost model to optimize the production of very large socs. IEEE Trans. Comput. 74 (4), 1278–1292. doi:10.1109/tc.2024.3521246 Jose Plathottam, S., Rzonca, A., Lakhnori, R., and Iloeje, C. O. (2023). A review of artificial intelligence applications in manufacturing operations. J. Adv. Manuf. Process. 5 (3), e10159. doi:10.1002/amp2.10159 Karim Kausik, A., Bin Rashid, A., Baki, R. F., and Maktum, Md M. J. (2025). Machine learning algorithms for manufacturing quality assurance: a systematic review of performance metrics and applications. Array 26, 100393. doi:10.1016/j.array.2025. 100393 Kazerouni, A., Ghavamzadeh, M., Abbasi-Yadkori, Y., and Van Roy, B. (2017). doi:10.1109/tcad.2016.2547904
- Conservative contextual linear bandits. Adv. Neural Inf. Process. Syst. 30, 3910–3919.
- Kim, F. (2010). “Chapter 1 - best practices in mission-assured, mission-critical, and safety-critical systems,” in Mission-Critical and Safety-Critical Systems Handbook. Editor K. Fowler (Boston: Newnes), 1–82.
- Kim, S. W., Kong, J.Ho, Lee, S. W., and Lee, S. (2022). Recent advances of artificial intelligence in manufacturing industrial sectors: a review. Int. Journal Precision Engineering Manufacturing 23 (1), 111–129. Kis, Á., Ancuti, C., and Ancuti, C. O. (2019). “Ats-pcb: an effective automated testing system for advanced driver assistance systems,” in 2019 International Symposium ELMAR, 215–218. doi:10.1007/s12541-021-00600-3
- Kulkarni, A., and Xu, C. (2021). A deep learning approach in optical inspection to detect hidden hardware trojans and secure cybersecurity in electronics manufacturing supply chains. Front. Mech. Eng. 7, 709924. Lackner, H., Thomas, M., Wartenberg, F., and Weißleder, S. (2014). “Model-based test design of product lines: raising test design to the product line level,” in 2014 IEEE Seventh International Conference on Software Testing, Verification and Validation, 51–60. doi:10.3389/fmech.2021.709924
- Lazebnik, T. (2023). Data-driven hospitals staff and resources allocation using agent-based simulation and deep reinforcement learning. Eng. Appl. Artif. Intell. 126, 106783. Lazebnik, T. (2026). The economical-ecological benefits of matching non-matching socks. arXiv Preprint arXiv:2602.18221. doi:10.48550/arXiv.2602.18221 Lazebnik, T., Golov, Y., Gurka, R., Harari, A., and Liberzon, A. (2024). Exploration–exploitation model of moth-inspired olfactory navigation. J. R. Soc. doi:10.1016/j.engappai.2023.106783
- Interface 21 (216), 20230746. Lazebnik, T., Aviv-Reuven, S., and Rosenfeld, A. (2025). Publishing instincts: an exploration-exploitation framework for studying academic publishing behavior and “home venues”. J. Inf. 19 (3), 101705. doi:10.1016/j.joi.2025.101705 Lervåg Synnes, E., and Welo, T. (2016). Bridging the gap between high and low-volume production through enhancement of integrative capabilities. Procedia Manuf. 5, 26–40. doi:10.1016/j.promfg.2016.08.006 Li, D., Xu, Ao, and Yu, X. (2023). “Optimized lightweight pcb real-time defect detection algorithm,” in 2023 IEEE 16th International Conference on Electronic Measurement & Instruments (ICEMI), 262–269. doi:10.1098/rsif.2023.0746
- Liang, X., and Zheng, D. (2022). “An industry example to reduce the test time by optimizing data extration method,” in 2022 China Semiconductor Technology International Conference (CSTIC), 1–4.
- Lima, J. A. P., and Vergilio, S. R. (2022). A multi-armed bandit approach for test case prioritization in continuous integration environments. IEEE Trans. Softw. Eng. 48 (2), 453–465. Liu, S., Peng, C., Chen, Z., Kan, Yu, Xiang, W., Li, J., et al. (2023). A learning-based context-aware quality test system in B5G-aided advanced manufacturing. IEEE Trans. doi:10.1109/tse.2020.2992428
- Industrial Inf. 19 (2), 1548–1558. Loecher, M. (2021). The perils of misspecified priors and optional stopping in multi-armed bandits. Front. Artificial Intelligence 4, 715690. doi:10.3389/frai.2021.715690 Mahajan, A., and Teneketzis, D. (2008). “Multi-armed bandit problems,” in Foundations and Applications of Sensor Management (Springer), 121–151. doi:10.1109/tii.2022.3169972
- Milor, L., and Sangiovanni-Vincentelli, A. L. (1994). Minimizing production test time to detect faults in analog circuits. IEEE Trans. Computer-Aided Des. Integr. Circuits Syst. 13 (6), 796–813. Mohamed, I., Mostafa, N. A., and El-assal, A. (2022). Quality monitoring in multistage manufacturing systems by using machine learning techniques. J. Intelligent Manuf. 33, 2471–2486. doi:10.1007/s10845-021-01792-1 Nie, H. (2024). Exploring the depths of multi-armed bandit algorithms: from theoretical foundations to modern applications. Appl. Comput. Eng. 68, 183–191. doi:10.54254/2755-2721/68/20241399 Pan, Y., Liang, H., Li, J., Huang, Z., Yi, M., and Lu, Y. (2025). Low test cost adaptive testing method for high yield ic products. Integration 103, 102401. doi:10.1016/j.vlsi. 2025.102401 Petkov, N., and Ivanova, M. (2024). Printed circuit board and printed circuit board assembly methods for testing and visual inspection: a review. Bull. Electr. Eng. Inf. 13, 2566–2585. doi:10.11591/eei.v13i4.7601 Pharmaceutical Inspection Co-operation Scheme (2009). Guide to Good Manufacturing Practice for Medicinal Products. Annexes PE, 009. Geneva, Switzerland: PIC/S Secretariat. doi:10.1109/43.285252
- Powell, D., Magnanini, M. C., Colledani, M., and Myklebust, O. (2022). Advancing zero defect manufacturing: a state-of-the-art perspective and future research directions. Comput. Industry 136, 103596. Quigley, J. M. (2024). Five approaches to product testing. IEEE Reliab. Mag. 1 (1), 30–36. doi:10.1109/mrl.2024.3356457 Rodrigues, N., Leitão, P., Foehr, M., Turrin, C., Pagani, A., and Decesari, R. (2013). “Adaptation of functional inspection test plan in a production line using a multi-agent system,” in 2013 IEEE International Symposium on Industrial Electronics, 1–6. doi:10.1016/j.compind.2021.103596
- Russo, D., Roy, B., Kazerouni, A., Osband, I., and Wen, Z. (2018). A Tutorial on Thompson Sampling.
- Saha, A., Ma, K., Amarnath, C., Qureshi, M., and Chatterjee, A. (2025). “Adaptive testing of compute-in-memory based cnns using probabilistic test acceptance limits,” in 2025 IEEE 31st International Symposium on On-Line Testing and Robust System Design (IOLTS), 1–7.
- Sankpal, K. A., and Metre, K. V. (2020). A review on data normalization techniques. Int. J. Eng. Res. 9 (6). Scheffler, M., Franzon, P. D., and Troster, G. (2004). A “defect level versus cost” system tradeoff for electronics manufacturing. IEEE Trans. Electron. Packag. Manuf. 27 (1), 67–76. doi:10.1109/tepm.2004.830513 Serban, M., Vagapov, Y., Chen, Z., Holme, R., and Lupin, S. (2014). Universal Platform for Pcb Functional Testing, 402–409. doi:10.17577/IJERTV9IS060915
- Slivkins, A. (2019). Introduction to multi-armed bandits. Found. Trends® Mach. Learn. 12 (1-2), 1–286. Song, T., Huang, Z., and Yan, A. (2022). Machine learning classification algorithm for vlsi test cost reduction. Integration 87, 40–48. doi:10.1016/j.vlsi.2022.06.005 Tahera, K., Earl, C., and Eckert, C. (2012). The Role of Testing in the Engineering Product Development Process. doi:10.1561/2200000068
- Tahera, K., Wynn, D., Earl, C., and Eckert, C. (2019). Testing in the incremental design and development of complex products. Res. Eng. Des. 30, 291–316. 018-0295-6 Tong, Z., Feng, J., and Liu, F. (2023). Understanding damage to and reparation of brand trust: a closer look at image congruity in the context of negative publicity. J. Prod. & Brand Manag. 32 (1), 157–170. doi:10.1108/jpbm-07-2021-3550 Tran, K. P. (2021). Artificial intelligence for smart manufacturing: methods and applications. Sensors 21 (08), 5584. doi:10.3390/s21165584 van den Burg, G. J. J., Nazábal, A., and Sutton, C. (2019). Wrangling messy csv files by detecting row and type patterns. Data Min. Knowl. Discov. 33, 11. doi:10.1007/s10618- 019-00646-y Vermorel, J., and Mohri, M. (2005). “Multi-armed bandit algorithms and empirical evaluation,” in European conference on machine learning (Springer), 437–448. doi:10.1007/s00163-
- Vijay Agrawal, A., Murthy Raju, K., Aparna, P., Sravya, G., Chandrashekhar, A., and Ramya, J. (2024). “Ai-driven test and measurement automation in electronics manufacturing,” in 2024 Ninth International Conference on Science Technology Engineering and Mathematics (ICONSTEM), 1–6.
- Voorakaranam, R., Cherubal, S., and Chatterjee, A. (2002). “A signature test framework for rapid production testing of rf circuits,” in Proceedings 2002 Design, Automation and Test in Europe Conference and Exhibition (IEEE), 186–191.
- Walston, J., Affan Badar, M., Kluse, C., and Rostom, R. (2025). A review analysis of cost of quality, cost of poor quality, and hidden quality cost. J. Technol. Stud. 50 (1–13), 1–13. Wang, X., and Yang, Z. (2025). “Optimal sampling and decision-making strategies in multi-stage manufacturing processes,” in 2025 10th International Conference on Cloud Computing and Big Data Analytics (ICCCBDA), 658–663. doi:10.21061/jts.435
- Wong, H. M., and Perumal, S. (2025). “Ai-driven model-retraining architecture to sustain operational accuracy in data-drifting environments,” in 2025 IEEE Symposium on Wireless Technology & Applications (ISWTA), 1–6.
- Wu, Y., Shariff, R., Lattimore, T., and Szepesvári, C. (2016). “Conservative bandits,” in Proceedings of the 33rd International Conference on Machine Learning (New York, NY: PMLR), 1254–1262.
- Yang, D.-H., Lee, H., and Kang, Y.-S. (2025). “Development of an ai model adaptation strategy based on multi-resolution concept drift detection: a case study in virtual metrology for smart manufacturing,” in 2025 IEEE Annual Congress on Artificial Intelligence of Things (AIoT), 335–341.
- Yilmaz, E., Ozev, S., and Butler, K. M. (2010). “Adaptive test flow for mixed-signal/rf circuits using learned information from device under test,” in 2010 IEEE International Test Conference, 1–10.
- Zahit Demiray, B., and Arslan, B. (2022). “Test cost-test quality modeling for adaptive test,” in 2022 IEEE International Conference on Automation, Quality and Testing, Robotics (AQTR), 1–5.
This page reproduces the article Peretz-Andersson et al. (2026), Frontiers in Mechnical Engineering, doi:10.3389/fmech.2026.1861443, with the permission of the publisher. Text, tables and figures were extracted from the PDF and the layout adapted for the web; the PDF is the version of record.
