Prueba Safe en un clic y descubre cómo mejoramos tu prevención de riesgos laborales

Prueba sin compromiso
artificial vision occupational safetyreal-time risk detectionAI accident preventionsmart industry camerassecurity monitoring

Cómo identificar riesgos en tiempo real con visión artificial

La visión artificial permite detectar riesgos en tiempo real antes de que se conviertan en accidentes. Así es como funciona aplicada a entornos industriales.

It's six in the morning, 1201. You get up to start your shift as a watchman at the Venetian Arsenal, so important that it left some of the first records of controls and protocols for safety at work.

That security falls on your shoulders, in an environment full of risks: warehouses full of gunpowder, primitive cranes, carpenters with an abundant ration working under pressure.

As you approach the gunpowder stores, you hear a metallic sound. Something doesn't fit. An operator, oblivious to the danger, hits the nails on a board with a hammer. A spark A flash. Silence.

You wake up. It's summer 2025. Grateful that many things have changed: access controls, preventive training, signage, personal protection... but the dream encourages us to ask ourselves, how can we continue to improve? What does the future of prevention hold for us, and how can we be part of it?

Today we talk about one of the most powerful and promising technologies in the prevention of occupational risks: artificial intelligence applied to artificial vision. Because now, By connecting a camera and a computer, you can create a system that observes your work environment 24 hours a day, detecting security breaches in real time, and provides you with objective data that supports your preventive work.

You no longer just watch: you demonstrate, communicate and raise awareness.

Sounds good right? Well let's go there.

First tests

One of the things I like most about all of this is that you don't need a sophisticated laboratory or big budgets for a start. Today we have access to free artificial intelligence models (yes,GRATIS!) available on the Internet.

A good example is Moondream, a well-known AI that allows you to ask questions in natural language about what appears in an image. You can try it yourself: upload an image of a person wearing a helmet and ask them if they are wearing their PPE correctly. You will see how he responds to you in a reasoned and conversational way.

We upload the image, click on the detect button, and ask it to search for a helmet. Sorry for the abuse of English but they work better that way, little trick.

This is called object detection, and is the basis for the simplest rules in computer vision.

But if we want to go further, it is not enough to just know what objects are there, but who does what with them. That is, identify the relationship between objects and people. We can try it from Moondream itself:

It's not bad at all, although we wanted him to take the whole person, and in almost all of them he has only selected the helmet

The most common (and best) way is to use the coordinates of the detections returned to us by the AI ​​model itself, which can be processed in a custom program: if a helmet is very close to a person's head, they are probably using it correctly, if a person is within coordinates marked as dangerous, an alert can be generated.

It looks good, right? Well this is just the beginning. Using detection techniques, we can succeed in a lot of interesting applications. Here are some of them:

Use cases

⚠ Dangerous areas and restricted zones

Perhaps one of the most efficient and effective cases to start with. The detection of people for these models is very good: shouldn't there be anyone in an area due to danger? Should there be at least 2 people in case one faints? Simply mark coordinates of safe/dangerous areas and we check which detection falls within.

👷‍♂ Detection of Personal Protective Equipment (PPE)

We already saw it in the example at the beginning, but we can continue adding endless equipment: glasses, vest, robe, high boots... As we explained, we detect people and objects, grouping by distance and overlap. If when grouping, someone is left without a security object, because they are too far away, we already have the offender.

🛠 Unsafe conditions (environment)

For example, presence of obstacles, distances to machinery that are too short, lack of collective protections or LOTOs. For the first case, the truth is that Moondream does a very good job, since it recognizes a large number of objects without us having to teach it anything, something that probably also helps with collective protections.

🔥 Dynamic risks

We see here cases such as the presence of fire, smoke, and one that I really like, the speed of forklifts. The truth is that those little bulls can be a bullet, and in the wrong hands a giant fork. For the first case, we let you try with the tools that we have seen being tested, while speed measurement has more details: it forces us at every moment to calculate the position of the object, measure its progress, and divide it by time, to obtain its speed. These techniques are usually related to techniques known as tracking. In the image, you can see how the objects leave a red trail, it is their tracking trail, in each of the positions that their center has been. With this we calculate its speed, and if it exceeds a certain threshold, alarm!

🧍‍♂ Human postures and actions

To detect falls or inappropriate positions, the YOLO has a magical ability: it allows you to detect the person's “skeleton”, to better understand if they have an abnormal position. This is often called pose estimation, and it generates a lot of interest in everything related to injuries and casualties. We attached a photo, but as before, you can try it yourself in this enlace.

That's about postures. The actions may have more substance, they deserve an almost special article. There are simple situations, such as detecting that someone smokes, that we can know by the cigarette or smoke. However, there are other situations that force us to follow and classify every moment of the people we check. This requires other technology, and can get a bit complicated. My little tip is: if the action has a recognizable moment, whether because an object or a posture is left, detect it and with that it will already be identified. If you need to analyze the entire video sequence, my best advice is to try Gemini 2.5

Continuous improvement

During testing you may have encountered errors in detection. This is normal. Generalist models are not trained for these types of situations, and now comes their true potential: they can be trained to specialize them, making them experts with human precision.

One of the most used for this, since training him is simple and fast, is the YOLO (also free, yeah). So popular that you can find it already prepared with the knowledge of this type of applications, as in Safe, which has seen tens of thousands of images to recognize a large number of situations and equipment, allowing you to enjoy maximum precision and reliability without wasting time. But you can also teach them from scratch. In a future article, we will explain how to do it, because it is simple, and I would say even fun (at least the first few hours).

And important note: perhaps you noticed that Moondream took several seconds to give you the results of your query. But, as the article is titled: How to identify risks in real time with computer vision. Get ready: YOLO detects image objects in about 50 milliseconds. This allows processing about 30 frames per second, which is common in most video surveillance cameras. What a Formula 1!

Next steps

Technology to monitor with a camera? Seen. Apply this system to review rules that improve the security of our environment? Also seen. Activate an alarm when an incident that we detect occurs? Generate a report with the results to present them in training? Graphs with personnel improvement trends? We have seen how to solve a key part: what is happening in the image, and we have to obtain our benefit, deciding what to do with that data based on its severity, or filter the information of greatest interest. In fact, we could even put an AI to cover the faces in the image, complying with the GDPR regulations. So we encourage you to continue learning about this technology, because the possibilities are exciting. 😀