Whenever the warehouse inventory check day came around twice a year, I would feel anxious for days beforehand.

Books are stacked in the warehouse on 3-tier racks. A mess...
Inside that dim, dusty 60-pyeong warehouse, climbing a wobbly ladder and straining my neck to look up at the 3-tier pallet racks was truly backbreaking work. The vinyl-wrapped bundles of books weren't even visible from below, and even when I shouted, "Boss, how many bundles is that yellow cover up there?" the boss, flipping through his ledger, would only blink behind his reading glasses. There was a time when I could just walk between the racks and instantly visualize it all in 3D in my head—"Ah, 400 English vocabulary books are stuck in that corner"—but now both the boss and I had equally faded photomemories.
We couldn't afford a logistics solution from a megacorporation costing tens of millions won, and scanning racks indoors with a drone using LIDAR was something I only saw in YouTube videos about other people's warehouses. Launching a drone in our cramped warehouse would just tear all the vinyl and crash into the pillars.
So I decided to go with a very simple and straightforward approach. I used a cheap Chinese multimodal LLM—where API costs are dirt cheap these days—as the backbone, and filled in the gaps with hands-on human work.

Since the rack specifications (2,590 mm wide × 1,000 mm deep) are fixed, we didn't need any fancy AI or expensive equipment to reverse-calculate "how many book bundle blocks fit in this space" by looking at perspective and depth in photos. With just 102 photos taken on my iPhone as I walked around and their timestamps, the 3D spatial layout according to the warehouse flow would fall into place.
But no matter how good the world becomes, there's a point where even if AI died and came back, it couldn't solve. When light reflects off the vinyl packaging and blurs the title, or when a corner of the cover sticks out, AI can't read it at all and stutters. But this is exactly where the usefulness of humans shines brilliantly.

When I display a thumbnail on the screen, the human eye catches it in 0.1 seconds. "Huh? See that yellow band around it? That must be the new one that came out." That fleeting visual cue that confuses AI even after tens of thousands of learning iterations—a human grasps it instantly from just one photo. As I look at the screen, playing like a puzzle game, I just say the name: "That's this book," and on the back end, the volume calculation algorithm computes the number of copies and plugs it into the ledger.
Half a day of physical labor climbing ladders and getting covered in dust has turned into a "10-minute hidden picture game" that I can finish in front of a monitor with a cup of coffee and a few clicks.
Without flashy cutting-edge equipment or drones, you lay down cheap technology in the right places and then put the finishing touches with what humans do best—their senses. In the end, solving the harsh problems of the field isn't about technological showiness; it's about this unique power of "the human eye and hand" that still shines so brightly.

I dream of AI and machines handing me cheap work to outsource, picking up digital scraps. Do I dream of lamb skewers? =3=3=3
▶ Original source: https://argo9.com/