Sigray is a manufacturer of scientific instruments - specifically those with advanced X-ray technology. As components like semiconductor designs become more complex, manually inspecting the output becomes a lofty challenge. For use cases like solder bump inspection, engineers require to review large stacks of 2D X-ray slices and visually scavenge for defects.
Sigray needed way to automate this process and include 3D X-ray datasets without adding more manual review. So, the GoML team built an AI powered X-ray defect detection and metrology system that does exactly that. Without any manual effort, the system converts 2D X-ray slices into a 3D volume, identifies and flags solder bump defects and automatically reports height/weight/width/volume at micron-level precision.
Built on GoML’s AI Matic platform, the system combines 3D segmentation, TensorRT-accelerated GPU inference and AWS serverless infrastructure to turn tedious manual inspection into an automated pipeline. Let’s learn how.
The challenge of massive 3D volumes and microscopic defects
Sigray needed to detect small structural defects such as bridges, cold joints, bulges and missing solder bumps inside very large CT scans - with a single volume measuring Z × 2940 × 2940 voxels using 16-bit data and input files commonly reaching 500 MB or more. Since the Z-axis changes with the physical size of each semiconductor, the team also had to handle variable volume depths while keeping the deep learning model within fixed input dimensions.
The data added another challenge because most voxels represented the background while some defects appeared only a few dozen times in a complete volume, with one dataset containing 40,000 normal instances compared with just 23 cold joints. The system also had to produce two outputs:
- Classify every voxel as background or one of five structural classes, then
- Convert those labels into physical measurements using the scanner's dimensional information from the slice.txt file.
Memory limits made the task harder, as standard sliding-window stitching can require around 15.7 GB of RAM, making it difficult to process these large volumes within the available hardware limits.
What made this AI powered semiconductor inspection system unique
Processing half-gigabyte X-ray volumes quickly while staying within strict GPU and memory limits required the team to solve three engineering challenges around 3D model processing, AWS infrastructure, and inference speed.
1. Memory-constrained 3D deep learning
Instead of loading an entire X-ray volume into memory, the team built a patch-based pipeline that processes only the data required for each prediction. The system reads volumes through memory mapping and divides them into fixed 64 × 192 × 192 micro-patches, keeping the model input size stable even when the Z-axis changes between semiconductor scans.
During inference, the system processes overlapping blocks one at a time, converts each block from a 6-channel float32 tensor to a 1-channel uint8 array, and clears it from memory before processing the next block, keeping peak system RAM at 4.8 GB.
2. AWS architecture for large files and GPU inference
The system uses FastAPI on AWS Lambda through Mangum for the web interface and secure S3 file transfers, while large files bypass Lambda's 6 MB payload limit and move directly from the browser to Amazon S3. This allows the system to handle input files of 500 MB or more without passing them through Lambda.
Once the files are ready, an Amazon SageMaker Async Endpoint running on ml.g 6.2xlarge instances handles 3D segmentation, with the NVIDIA L4 GPU keeping memory usage below 22 GB during inference. Lambda, S3, and SageMaker Async can also scale down when no jobs are running, so GPU resources do not remain active between scans.
3. TensorRT acceleration and engine caching
The team converted the ONNX model into a TensorRT FP16 engine to reduce inference time while maintaining accuracy. Since compiling the engine can take 5 to 15 minutes, the system stores the compiled engine in S3 and repackages it with the deployment so later restarts and scaling events can begin with a ready-to-use engine, bringing warm-start time down to seconds.
The novelty of how GoML approached this challenge
The GoML team delivers best-in-segment production AI systems in record time as it is powered by the rapidly reusable components and infrastructure of the AI Matic platform. We built Sigray's AI X-Ray defect detection and metrology system, giving the engineering team solution bundles for secure data ingestion, model packaging, serverless deployment and inference orchestration. GoML’s Ninja FDEs were then used to adapt the code to process volumetric 3D X-ray data, run TensorRT-optimized GPU inference and also calculate micron-level metrology while keeping the results layers separate.
We specifically used the AI Data Analyst solution blueprint as a starting point because the system needed to turn raw X-ray slices into structured measurements. Before finalizing the design, testing was done with two approaches within a limit of 12.7 GB RAM and 15 GB VRAM: SwinUNETR-based 3D segmentation and PointPillars-based 3D object detection. The project covered four requirements, all of which were delivered in record time:
1) Defect classification
2) Metrology measurement
3) 3D image model integration, and
4) A demo interface
The workflow starts when engineers upload TIFF slices or a ZIP file with the slice.txt metadata file directly to S3, after which the system creates a unique job ID. Validation checks the file types, slice numbering, dimensions plus metadata before processing, so incomplete scans with missing slices do not reach the GPU. The system then stacks the 2D slices into a 3D volume and applies z-score normalization to the 16-bit input.
The MONAI SegResNetDS model processes overlapping patches through the TensorRT engine and stitches the predictions back to the original volume size. Connected component analysis then separates individual solder bumps from the predicted mask and calculates their height, width, and volume in microns. The SageMaker container packages the predicted mask into a ZIP file and stores it in S3, while engineers can upload multiple scans, save their job IDs also run inference later. Completed results are available through short-lived pre-signed GET URLs.
"For this build we did for Sigray, accuracy had to go hand in hand with performance which is a high-stakes challenge. All constraints were considered like the 3D X-ray data, variable scan volumes and limited GPU memory. The outcome is a system that will prove to be a gamechanger for Sigray" - Prashanna Hanumantha Rao, Co-Founder and CTO, GoML
How the system automated semiconductor inspection for Sigray
The system brings X-ray uploads, validation, defect detection, metrology and result delivery into one workflow - drastically reducing the need for engineers to move between CT slices and separate measurement steps. Instead of reviewing stacks manually and recording bump measurements by hand, engineers can upload a scan and let the system process the complete volume in one go.
Validation checks each upload for missing slices, missing metadata, unsupported file types including mismatched dimensions before GPU processing begins. This prevents incomplete or incorrect datasets from consuming inference resources.
The metrology stage also removes manual measurement work by extracting dimensions directly from the 3D voxel masks. The system calculates the height, width with volume of each solder bump in microns using the scanner's physical metadata.
The AWS setup also limits unnecessary compute usage. Lambda, S3, and SageMaker Async can scale down when no jobs are running, while GPU resources are used when a scan enters inference. TensorRT engines are compiled once and then reused from cached files, avoiding another 5 to 15 minutes of compilation after every restart.
The frontend does not handle large raw files directly. Uploads move through short-lived pre-signed PUT URLs, while heavy processing and ZIP creation take place in the SageMaker container. This keeps large data and processing workloads away from the web layer.
Security controls apply to each file transfer and AWS operation. Uploads use pre-signed PUT URLs, downloads use pre-signed GET URLs, and IAM roles control AWS access. User actions are limited to their generated job IDs, while each job ID connects to upload, inference payload and result mask stored in S3.
Winning outcomes
Here are some of the winning metrics that the AI-powered system enabled for Sigray:
- Analysed accurately in under 2 minutes: A full 3D X-ray volume can be analyzed in under two minutes, with production runs reaching about one minute.
- 70% lower peak system RAM: Inference uses about 4.8 GB of peak system RAM compared with roughly 15.7 GB for standard sliding-window stitching.
- Seconds to warm-start: Cached TensorRT engines allow the endpoint to start in seconds instead of repeating the 5-15 minute engine compilation process.
How GoML delivered the system
AI Matic was the foundation stone for the system which enabled large file uploads, secure access, GPU inference and also model flexibility. We focussed on Sigray's exact needs and built a bespoke system that accounts for all logistical constraints.

Explore more cutting-edge production AI systems in GoML's case study section, or connect with our experts to get started on your AI roadmap.





