Skip to main content

Machine Vision in Material Handling: What It Does and How It Works

Machine vision gives a material handling system a way to see the part it is about to move: where the part is, which way it faces, whether it is the right part, and whether it is fit to pass to the next station. UTEC Industrial designs, engineers, machines, fabricates, and installs custom material handling systems for aerospace and heavy industry from its Spokane Valley, WA facility, integrating Allen-Bradley PLC and motion control with in-house CNC machining, heat treating, and stress relief. This article explains what a vision system does on a plant floor, how its imaging chain of lighting, lens, sensor, interface, and software produces a usable result, how a camera is calibrated to guide a robot, where deep learning helps, and why a production camera is not a safety device. Vision sits near the end of the build chain, design → engineering → parts machining → fabrication → assembly → weld fatigue → stress relief → drives → controls → tuning → monitoring, but it depends on every mechanical link before it.

What does machine vision actually do in a material handling system?​

A vision system turns an image into a decision or a coordinate that the handling equipment can act on. The Hornberg handbook's manufacturing chapter, in its section on 3D systems, lists applications under six headings. Each heading maps to a handling use:

  • Identification. Reading a code, a part number, or a type feature so the line routes the part correctly, for example confirming which of several casting patterns has arrived on a foundry conveyor.
  • Completeness check. Confirming that every fastener, clip, or insert is present before an aerospace subassembly moves to the next station.
  • Object and pose recognition. Finding where a part lies and how it is oriented, so a robot or transfer mechanism can pick it without a hard fixture.
  • Shape and dimension measurement. Gauging a profile, a hole pattern, or a board's width and thickness in line.
  • Surface inspection. Detecting cracks, stains, knots, or coating defects.
  • Robotics. Supplying the robot with a corrected pick or place position.

The handbook also separates discrete unit production from continuous flow, and job-shop production from mass production. That distinction matters for handling: a sawmill board line or a strip line is continuous flow, while an airframe fitting moving between stations is a discrete unit imaged when it stops or passes a trigger point. The handbook's temporal-interface section treats discrete motion production, continuous motion production, and line-scan processing separately, so the production type shapes when and how the image is taken (Hornberg 2017, Ch. 10, §10.2 Application Categories, §10.8 Temporal Interfaces and §10.10.3 Applications).

Where does a vision system sit in the material handling build chain?​

Vision is part of the intelligence layer, the controls, tuning, and monitoring at the end of the chain, but its accuracy is set upstream. The handbook's integration section lists the mechanical interfaces a vision system depends on: dimensions and fixation, working distances, position tolerances, forced constraints, additional sensor requirements, and additional motion requirements. Each maps onto an earlier link in the build chain:

  • Design and engineering fix the camera's working distance and field of view and decide whether the part is presented in a known position or found anywhere in a bin.
  • Parts machining sets the repeatability of the fixtures, nests, and datum stops that present a part to the camera. A nest that varies by more than the camera's measurement tolerance makes every downstream measurement uncertain.
  • Fabrication, weld fatigue, and stress relief determine whether the frame carrying the camera and the part stays dimensionally stable. A mounting frame that moves as residual weld stress relaxes, or that vibrates with a nearby drive, shifts the camera relative to the calibration it was commissioned with.
  • Drives, controls, and tuning decide when the image is taken, how the result reaches the PLC, and how the robot or transfer car uses it.

UTEC Industrial machines fixtures and mounting interfaces to tolerances of ±0.001 in, which is the kind of mechanical repeatability a vision measurement relies on. The handbook covers reproducibility and gauge capability within that same mechanical-interfaces section (Hornberg 2017, Ch. 10, §10.5 Mechanical Interfaces, §10.5.8 Reproducibility and §10.5.9 Gauge Capability).

What are the parts of a machine vision imaging chain?​

Every vision system is a chain of five parts, and the result is only as good as the weakest link:

  1. Illumination, which creates the contrast the software needs.
  2. Optics, the lens that forms the image, sets the field of view, and limits depth of field.
  3. The camera and image sensor, which convert light to pixel values with a characteristic noise, sensitivity, and dynamic range.
  4. The camera-computer interface, such as GigE Vision or Camera Link, which moves image data at the pixel rate the application needs.
  5. Image processing software, which segments, measures, matches, or classifies the image.

The Hornberg handbook follows broadly the same order, with chapters on lighting, optical systems, camera calibration, camera systems, smart cameras, camera-computer interfaces, and machine vision algorithms. Its algorithms chapter covers thresholding and segmentation, edge extraction, fitting of lines, circles, and ellipses, template matching, stereo reconstruction, and optical character recognition. For sensors, the handbook covers CCD and CMOS technology, CMOS shutter concepts, sensor sizes, a section on qualifying cameras and noise measurement, and a section on camera noise that includes fixed pattern noise. A common failure mode is to buy the camera first and treat lighting and optics as accessories; the chain is designed from the part and its contrast outward (Hornberg 2017, Ch. 3 to Ch. 9).

How are camera and image sensor specifications compared?​

Camera data sheets are only comparable when they are measured the same way. EMVA 1288 is the European Machine Vision Association's standard for the measurement and presentation of specifications for machine vision sensors and cameras; its stated aim is a unified method to measure, compute, and present specification parameters for cameras and image sensors used in machine vision. Release 4.0 became effective in June 2021. The earlier Release 3.1, from December 2016, used a simple linear model and was therefore limited to cameras with a linear response and without any pre-processing; Release 4.0 takes into account the rapid development of camera technology since then.

For a specifying engineer this has a practical consequence. Two cameras quoted with the same resolution can differ widely in how much signal they deliver per unit of light and how much noise they add, and that difference decides whether a dark casting, a wet log, or a machined aluminum surface under strong reflection can be imaged reliably at the line speed. Asking for EMVA 1288 data, and noting which release the data was measured to, is the way to compare sensors on measured performance rather than on marketing figures (EMVA 1288 Release 4.0-2021).

How is camera resolution matched to the part and the field of view?​

Resolution is worked backwards from the smallest feature that has to be seen. The handbook's system-design chapter walks through camera type, field of view, camera sensor resolution, spatial resolution, measurement accuracy, and a separate calculation for line-scan cameras, in that order.

An illustrative calculation, with assumed numbers rather than values taken from any source, shows the idea. Assume a lumber line needs to see a board 48 in wide, and the camera sensor is 2,448 pixels across:

  • Spatial resolution = field of view ÷ pixels = 48 in ÷ 2,448 px = 0.0196 in per pixel.
  • A defect 0.10 in across then covers 0.10 ÷ 0.0196 = about 5 pixels.

Assumptions: the full 48 in width is imaged by one camera with no margin, the lens adds no distortion, and the part stays at the focused working distance. In practice a margin is added for part position tolerance, and if 5 pixels is too few for the software to classify the defect, the options are a higher-resolution sensor, a second camera, or a narrower field of view per camera. For a continuous-flow line, the handbook treats line-scan resolution in its own section; along the direction of travel it also depends on line rate and conveyor speed. In the handbook's sequence, resolution is followed by the choice of camera, frame grabber, and hardware platform, and then by lens design, with sections on focal length, lens diameter and sensor size, and sensor resolution and lens quality (Hornberg 2017, Ch. 2, Designing a Machine Vision System, pp. 36–43).

Why is lighting the first design decision in a vision system?​

The software can only measure contrast that the lighting creates. The handbook's system-design chapter states the lighting concept as maximizing contrast, and its lighting chapter devotes more than 100 pages to it, covering:

  • Lighting techniques, including diffuse and directed bright field, telecentric bright field, structured bright field, dark field, and transmitted (backlit) arrangements.
  • Light color and part color, including monochromatic, white, infrared, and ultraviolet light, and polarized light.
  • Light filters, including daylight suppression filters, neutral density filters, and polarization filters.
  • Lighting control, including brightness control and static, pulsed, and flash operation.
  • Suppression of ambient and extraneous light.
  • Lifetime, aging, and drift of light sources.

These topics map to common failure modes on a plant floor. Sunlight through an open bay door or a skylight changes the image between morning and afternoon shifts. A light source that dims with age slowly pushes a threshold-based inspection toward false rejects. Radiant glow from a hot billet in a steel or aluminum mill, or steam in a pulp mill, adds light the system was not designed for. Flash illumination synchronized to the camera exposure, with filters matched to the light source, is a common defense; the handbook treats static, pulsed, and flash operation and the suppression of ambient light in its lighting-control section (Hornberg 2017, Ch. 2 p. 44 and Ch. 3 pp. 63–177).

When does material handling need 3D vision instead of 2D?​

A 2D camera measures position and rotation in a plane; it cannot tell how high a part sits or how it is tilted. That is enough when parts arrive flat on a conveyor at a known height. It is not enough when parts are stacked, piled in a bin, or presented at varying heights, which is common for castings, forgings, and machined blanks delivered in returnable containers.

The handbook's 3D section separates 2.5D from 3D data and covers point clouds and registration. It divides 3D data acquisition into passive and active methods, and its algorithms chapter covers stereo reconstruction, including stereo geometry and stereo matching.

Robot makers now package 3D detection for random bin picking. FANUC's iRVision 3D Model Detection Function automatically generates the settings needed to detect a part in various positions from the part's 3D CAD data, for bin picking in which a robot picks one of many parts placed randomly in a returnable container using a 3D vision sensor. FANUC notes that the settings for each possible part position previously had to be made and registered manually, one by one. The detection processing runs on a PANEL iH Pro connected to the robot controller by an Ethernet cable (Hornberg 2017, Ch. 10, §10.10 3D Systems; FANUC Corporation 2021, iRVision 3D Model Detection Function).

How is a camera calibrated so a robot can use what it sees?​

A robot cannot use a pixel coordinate; it needs a position in its own coordinate frame. Getting from one to the other takes two calibrations:

  • Camera (intrinsic) calibration models the lens and sensor so pixel positions can be converted to geometric rays. Zhang's planar-pattern method requires the camera to observe a flat pattern shown at a few, at least two, different orientations. Either the camera or the pattern can be moved freely, and the motion need not be known. The method models radial lens distortion and uses a closed-form solution followed by a nonlinear refinement based on the maximum likelihood criterion.
  • Hand-eye calibration finds the position and orientation of the camera relative to the robot. Tsai and Lenz described computing the camera pose relative to the last joint of the robot in an eye-on-hand configuration from a series of automatically planned robot movements. In their system, each move took a total of 90 ms to grab an image, extract image features, and perform the camera's extrinsic calibration, and the hand-eye computation took about 100 + 64N arithmetic operations, where N is the number of stations. They also analyzed the critical factors influencing accuracy.

The handbook's robot-guidance case study lists calibration and communication as its two key points. A common failure mode is a camera bracket that is bumped, loosened, or replaced after commissioning: the images still look correct, but every pick is offset until hand-eye calibration is repeated. UTEC Industrial integrates FANUC robotic cells, including vision, with a FANUC design and engineering partner (Zhang 2000, IEEE TPAMI 22-11; Tsai and Lenz 1989, IEEE Transactions on Robotics and Automation 5-3; Hornberg 2017, Ch. 10, §10.11.5 Robot Guidance).

How well does deep learning handle bin picking?​

Deep learning is most useful where a part has too little texture or too much variation for classic template matching. Jiang and co-authors built a bin-picking system for textureless boxes that uses only depth images from an Intel RealSense SR300 camera mounted on the suction hand in an eye-in-hand arrangement. A deep convolutional neural network trained on 15,000 annotated depth images, generated synthetically in a physics simulator, predicts grasp points without first segmenting each object, and predicts the vacuum-cup grasp pattern for the two-cup hand. The reported results were:

  • A 97.5% pick success rate at speeds exceeding 1,000 pieces per hour, with a 7-degree-of-freedom robot on randomly posed boxes in clutter.
  • A grasp planning time of 0.034 s.
  • 79.16% of predicted grasp points within 30 pixels of the surface center, compared with 58.88% for Dex-Net 4.0.

The failure modes are just as instructive. Picks failed when the target box was adjacent to another box and the robot targeted the wrong box, and when the surface area was close to the size of the vacuum cup. The authors state that the system handles only planar-faced objects with rectangular contours. The authors tie that limit to a training set generated only from boxes. The same bounding lesson applies to inspection more broadly: a network trained on a defined set of surface or dimensional defect classes only recognizes those classes, so a new defect class is a new training-and-validation job, not a configuration change. The research was set in warehouse logistics, not a heavy plant, but the lesson transfers: a trained model is bounded by the part shapes it was trained on, and a new part family is a new validation job (Jiang et al. 2020, Sensors 20-3, 706, §7 and §8.4).

How does a vision system connect to the PLC and the handling controls?​

The vision result is only useful if it reaches the controller at the right moment and in a form the controller can trust. The handbook's interface sections cover digital I/O, field buses, serial interfaces, standard Ethernet TCP/IP, OPC UA, and Ethernet-based real-time field buses, plus timing for discrete-motion, continuous-motion, and line-scan production. On an Allen-Bradley system the design details come from Rockwell Automation's controller documentation:

  • Event-driven logic. Logix 5000 controllers support continuous, periodic, and event tasks. Event tasks can be triggered by a consumed tag, an EVENT instruction, a change in module input data, or a motion event, so a camera result can start a sequence without waiting for the next scan of a continuous task.
  • Handshaking. Produced tags transmit continually at the requested packet interval, so it can be difficult to know when new data has arrived. Rockwell Automation recommends embedding a bit or counter in the produced tag to flag new data, and a return handshake so the producer knows the consumer has received and processed it. A common failure mode without that handshake is a PLC acting twice on one stale pick coordinate.
  • Data integrity. The CPS instruction buffers produced and consumed data and provides data integrity for data structures larger than 32 bits, which matters for a multi-value pose result.
  • Sensing around the camera. Part-present sensors, encoders on the conveyor or transfer car, and triggers from the drive tell the camera when to image and tell the PLC where the imaged part has moved to.

UTEC Industrial, a Rockwell Automation Recognized System Integrator, programs Allen-Bradley ControlLogix and CompactLogix controllers over EtherNet/IP for the handling systems it builds (Hornberg 2017, Ch. 10, §10.7 Information Interfaces and §10.8 Temporal Interfaces; Rockwell Automation 1756-RM094N-EN-P-2025).

Can a machine vision camera be used as a safety device?​

Not unless it is built and tested as safety equipment. IEC 61496-1:2020 sets general requirements for the design, construction, and testing of non-contact electro-sensitive protective equipment (ESPE) designed specifically to detect persons or parts of persons as part of a safety-related system. It does not deal with requirements for ESPE functions not related to the protection of persons, such as using sensing unit data for navigation, and it notes that applications such as protecting machinery or products from mechanical damage can call for different requirements. A production camera that finds parts is not designed specifically to detect persons, so it is not a substitute for a safety light curtain or laser scanner that meets the standard.

Where a light curtain or scanner guards a vision-guided robot cell, ISO 13855:2024 covers positioning the safeguard with respect to the approach of the human body. The safety logic belongs in a safety controller: in a GuardLogix system, only the safety task, not standard tasks, can be used for safety functions. Rockwell Automation rates a GuardLogix 5580 primary controller with a safety partner up to SIL 3 and PLe (Cat. 4), and one without a partner up to SIL 2 and PLd (Cat. 3). The design of these safety-related control parts falls under ISO 13849-1:2023, the underlying risk assessment under ISO 12100:2010, and the machine's electrical equipment under IEC 60204-1:2016. A common failure mode is a vision "presence check" quietly promoted to a personnel interlock (IEC 61496-1:2020; ISO 13855:2024; Rockwell Automation 1756-RM012J-EN-P-2025; ISO 13849-1:2023; ISO 12100:2010; IEC 60204-1:2016).

What must happen before someone cleans or adjusts a camera inside a cell?​

Lenses and lighting windows in a foundry, a sawmill, or a biomass plant collect dust and need regular cleaning, and every cleaning trip puts a person inside the robot or conveyor envelope. OSHA's lockout/tagout standard applies to that servicing work. It defines push buttons, selector switches, and other control-circuit-type devices as not being energy-isolating devices, so stopping the robot from the HMI or opening the camera's software does not isolate the cell. Under 1910.147(d)(5)(i), after lockout devices are applied, all potentially hazardous stored or residual energy must be relieved, disconnected, restrained, and otherwise rendered safe. In a vision-guided cell that includes a robot arm holding a part under gravity and pneumatic pressure in a gripper or vacuum system.

The handbook's remote-maintenance section for vision systems carries a safety precaution headed "No Movements." A good design places cameras and lights where they can be reached from outside the guarded zone, or provides windows that can be cleaned without entry, so routine cleaning does not require a full lockout (OSHA 29 CFR 1910.147-1989; Hornberg 2017, Ch. 10, §10.9.3.1 Safety Precaution: No Movements).

How is a vision system tuned, validated, and monitored after startup?​

A vision system commissioned on a clean day with new lights is not proven until it has run through real production variation. The handbook's project-realization section covers development and installation, a test run and acceptance test, and training and documentation, and its calibration chapter includes verification of calibration results. The work after startup covers:

  • Acceptance on production parts. The test run uses parts across the real range of color, surface finish, temperature, and position, not a single golden sample.
  • Gauge capability. Where the vision system measures, its repeatability is checked the same way as any other gauge, which the handbook covers under reproducibility and gauge capability.
  • Drift monitoring. Light source aging and drift, lens contamination, and camera mount movement all show up as a slow change in image brightness, contrast, or match scores. Trending those values flags a problem before the reject rate climbs.
  • Traceability and result data. Storing results and images per part supports root-cause work when a defect escapes.
  • Recalibration triggers. Any bracket, lens, or robot tool change is a trigger to verify calibration before production resumes.

UTEC Industrial performs factory acceptance testing and on-site commissioning, so acceptance criteria for a vision-equipped handling system can be written into the purchase order and demonstrated before handover (Hornberg 2017, Ch. 2 p. 49, Ch. 3 p. 88, Ch. 5 p. 308 and Ch. 10 §10.7.2).

What should a specifying engineer define before buying a vision system?​

The handbook's specification checklist gives a structure for a request for quotation. A complete vision specification for a handling system defines:

  • Task and benefit: what decision the system makes and what it saves, such as fewer mis-picks or fewer escaped defects.
  • Parts: the part types, their variation in color, finish, and geometry, and how each is presented to the camera.
  • Performance requirements: measurement accuracy and time performance, meaning the cycle time or line speed the result must keep up with.
  • Information interfaces: what data goes to the PLC or robot, in what format, and over which network.
  • Installation space and environment: mounting space, working distance, ambient light, temperature, dust, moisture, and vibration.

Two common failure modes follow when this list is skipped. The first is a system that works on the bench but misses cycle time on the line, because time performance was never specified. The second is a system that meets accuracy with a perfect sample but fails on the real spread of parts, because part variation was never described. Writing both into the specification lets the supplier size the lighting, sensor, and processing before anything is built (Hornberg 2017, Ch. 2, Specifying a Machine Vision System, pp. 32–35).

Related Articles

References​

  • Hornberg, A. (Ed.). Handbook of Machine and Computer Vision: The Guide for Developers and Users, 2nd ed. Wiley-VCH, 2017. ISBN 9783527413393
  • EMVA 1288 Release 4.0-2021: Standard for Measurement and Presentation of Specifications for Machine Vision Sensors and Cameras. European Machine Vision Association, 2021.
  • FANUC Corporation (2021). "ROBOT New Function: iRVision 3D Model Detection Function." FANUC News, 2021.
  • Zhang, Z. (2000). "A flexible new technique for camera calibration." IEEE Transactions on Pattern Analysis and Machine Intelligence, 22(11), 1330-1334. DOI 10.1109/34.888718
  • Tsai, R. Y., & Lenz, R. K. (1989). "A new technique for fully autonomous and efficient 3D robotics hand/eye calibration." IEEE Transactions on Robotics and Automation, 5(3), 345-358. DOI 10.1109/70.34770
  • Jiang, P., Ishihara, Y., Sugiyama, N., Oaki, J., Tokura, S., Sugahara, A., & Ogawa, A. (2020). "Depth Image-Based Deep Learning of Grasp Planning for Textureless Planar-Faced Objects in Vision-Guided Robotic Bin-Picking." Sensors, 20(3), 706. DOI 10.3390/s20030706
  • Rockwell Automation 1756-RM094N-EN-P-2025: Logix 5000 Controllers Design Considerations. Rockwell Automation, 2025.
  • IEC 61496-1:2020: Safety of machinery — Electro-sensitive protective equipment — Part 1: General requirements and tests. IEC, 2020 (Ed.4).
  • ISO 13855:2024: Safety of machinery — Positioning of safeguards with respect to the approach of the human body. ISO, 2024.
  • Rockwell Automation 1756-RM012J-EN-P-2025: GuardLogix 5580 and Compact GuardLogix 5380 Controllers Safety Reference Manual. Rockwell Automation, 2025.
  • ISO 13849-1:2023: Safety of machinery — Safety-related parts of control systems — Part 1: General principles for design. International Organization for Standardization, 2023.
  • ISO 12100:2010: Safety of machinery — General principles for design — Risk assessment and risk reduction. ISO, 2010.
  • IEC 60204-1:2016 (Ed. 6.0): Safety of Machinery -- Electrical Equipment of Machines -- Part 1: General Requirements. International Electrotechnical Commission, 2016.
  • OSHA 29 CFR 1910.147-1989: The Control of Hazardous Energy (Lockout/Tagout). Occupational Safety and Health Administration, 1989.

Ready to Discuss a Material Handling System?​

UTEC Industrial designs, engineers, machines, fabricates, and installs custom material handling systems for heavy industry, from the stress-relieved structure and drives to the Allen-Bradley PLC controls, tuning, and monitoring that run them, at its Spokane Valley, WA facility. Send UTEC the application, loads, and duty cycle to start a system review.

Request a Quote →

Questions? Call (509) 922-1832 or email sales@utec.co