We build detection, recognition and tracking that run on live camera feeds around the clock — not on a test set. First production build in 3-6 weeks.
Computer vision development services cover the whole path from a business rule to a model running in production: data collection and labelling, model selection and training, the inference pipeline, and the integration that puts results where your team already works. Fora Soft has built 250+ software products since 2005, and our computer vision runs on live video rather than stored images.
Most vendors stop at the model. On live video the model is the smaller half of the job. A detector that scores 99% on a benchmark can still miss a person at dusk on camera 34, because the hard part is the stream feeding it: decode, frame sampling, GPU budget, back-pressure when a link drops.
One camera or four hundred — the architecture is the same, and so is where we start.
You have cameras and a platform. You want the footage to answer questions instead of just recording.
Your VMS detects motion and nothing else. You need rules your own operators can change.
The model is the product. You need it accurate enough to sell, and cheap enough per stream to scale.
Computer vision demos succeed on curated clips and fail on live feeds. Four things break them: lighting the training set never saw, the cost of running inference on every frame, models drifting as cameras and scenes change, and storage bills nobody modelled. Each is an engineering decision, not a model choice.
A model trained on daylight footage loses people at dusk. We label from your cameras, in your conditions, and hold back a night-and-rain set the model never sees during training.
Running detection on every frame of 40 cameras is the line item that kills the business case. Frame sampling, region-of-interest cropping and cascades cut GPU spend by an order of magnitude. We size this before the contract, not after.
Cameras get moved, lenses fog, seasons change, accuracy slides — usually unnoticed. We ship the monitoring alongside the model: per-camera confidence trends, and a retraining loop with a labelled queue.
Video storage is the cost that compounds. Event-only retention, edge inference, cold tiers — we model the three-year bill before writing the ingest.
Six things a camera can be taught to do. Each one has shipped to production at least once, and the number beside it is from that deployment.
Object detection and tracking on live streams: people, vehicles, equipment, PPE, products. YOLO-family and DETR detectors, DeepSORT/ByteTrack for identity across frames, tuned for your camera placement rather than a public benchmark.
Face recognition with anti-spoofing. On Mindbox we shipped 99.5%+ recognition accuracy with resistance to photo and video replay attacks.
Licence plates, labels, meters, documents in frame. The ANPR module we built for Mindbox reads 500,000+ vehicles a day across India at roughly 95% accuracy.
Anomaly and event detection: intrusion, fire, crowding, falls, hard-hat violations, cashier fraud. Live Eye matches POS transactions against the video and alerts the owner within 30 seconds.
Counts, dwell time, heatmaps, queue length, process timings — piped into the dashboards and BI your team already opens, not a second console nobody logs into.
Forensic search over recorded footage by object, attribute or spoken word. On VALT, transcription and word search let 50,000+ users jump to the moment inside a recording and export it as a PDF.
Five steps, and what you get at the end of each one.
One call. A few days later: structured requirements, the recommended models and stack, and a block-level estimate. Free.
We pull footage from your cameras, label a first set, and stand up a baseline model. You see real numbers on your own video before committing to the full build.
Ingest, decode, sampling, inference, storage. The cost per stream is fixed here, and it is the number that decides whether the project makes sense.
Engineers orchestrate AI agents, review every change by hand, and write the hard parts themselves. Demos monthly or whenever you ask.
Rollout to production, per-camera accuracy monitoring, and a retraining loop. Accuracy that holds in month nine is the deliverable, not accuracy in week one.
The steps above are illustrative. Every project follows its own path, and we’ll agree on the specific plan before work begins.
The right approach depends on your product. We can use a cloud vision API, integrate an off-the-shelf model, or build a custom one around your data and requirements. We’ll help you weigh accuracy, cost, deployment, and ownership — and build the solution that fits.
Your task is common — reading text, generic object labels, generic face match — and a per-call price you can live with. Off-the-shelf gets you there in days. Take it.
The rule is yours: PPE on this site, this SKU in this aisle, this behaviour at this door. No vendor has your footage, so no vendor's model has seen your conditions. Everything you build stays yours, at any size — from the first camera.
We have done both, and we have talked clients out of the second one. Free architecture review, and you keep the answer either way.
Get a free architecture reviewA first production build — one use case, your cameras, model plus pipeline plus a dashboard your team can actually use — starts at $22,500: 450 engineering hours over 3-6 weeks.
Indicative pricing. The price and timeline above are starting points for this type of build. Every project is different, and we’ll confirm the scope, estimate, and delivery plan with you before work begins.
How most of our clients work
When you need one number signed off
Scope is written, price is fixed. Best when the use case is one and clear.
Scope moves as accuracy data comes in. Most CV projects end up here, and we say so before you sign.
Architecture review, model audit, cost-per-stream sizing. Days, not months. Often the honest answer is "don't build this".
Our engineers inside your process. If you want to hire CV engineers into your own team instead, we wrote how to do it.
How to hire computer vision developers →
Send the use case and a few minutes of footage. You get back the models worth trying, the cost per stream, and whether an off-the-shelf API would do.
Get it free →
Already have a model that underperforms in production? We look at the data, the pipeline and the thresholds, and tell you which of the three is the problem.
Get it free →
After one discovery call: structured requirements, recommended technologies, block-level estimate.
Get it free →
A specialist review of your video or streaming product covering latency, media server architecture, WebRTC, playback reliability, real-time chat, and scalability. Every finding is specific, located, and fixable. Delivered within a week.
Get it free →Five questions worth asking every vendor on your shortlist, including us. Our answers are underneath each one.
Our answer.Mindbox runs 99.5%+ face recognition and reads 500,000+ licence plates a day in production. Live Eye processes 2,000+ real-time interactions daily. Neither is a benchmark score.
Our answer.We have built video infrastructure since 2005 — WebRTC, RTSP, media servers, 250+ products. When a CV project breaks, it usually breaks in the stream, and that is the part most CV shops outsource.
Our answer.We size inference and storage during the free architecture review. If the numbers do not work, you find out before you have paid us anything.
Our answer.Mindbox, Live Eye Surveillance, EyeBuild, VALT — named, live, with metrics. Four of the six pages ranking for this search name no clients at all, or name clients with no numbers. Clutch named Fora Soft a Global Spring 2024 leader in computer vision.
Our answer.You do. All of it, including the labelled data. Written into the contract.
Accuracy, cost, limits, integration, data and ownership.
Ask an engineerOn production systems we have shipped, 99.5%+ for face recognition and around 95% for licence plates at 500,000 vehicles a day. Your number depends on camera placement, lighting and how rare the event is. We measure it on your footage during the free architecture review, before you commit.
A first production build — one use case, your cameras, model plus pipeline plus dashboard — starts at $22,500: 450 engineering hours over 10–14 weeks. Cameras, event rarity, accuracy target and retention window move it. Fora Soft gives a block-level estimate free after one discovery call.
Ask three questions: has their model run on live video or only on a dataset, will they size cost per stream before the contract, and who owns the weights afterwards. Fora Soft has built video infrastructure since 2005 across 250+ projects, and you own everything we train.
Yes, and that is the point of the free architecture review. Some problems are better solved with a sensor, a barcode or a process change. We have talked clients out of building. You keep the answer either way.
It reads what a camera can see. Intent, ownership and context it infers badly. Heavy occlusion, extreme low light and rare events with few examples all cut accuracy hard. Anyone promising 99% on every use case has not run one in production for a year.
Usually yes. We work over RTSP and ONVIF with standard IP cameras, and we have integrated with POS for transaction-to-video matching on Live Eye and built an IVMS with a REST API and SDK on Mindbox.
Yes. EyeBuild runs on solar-powered cameras over 4G/5G with no internet connection and a 14-day battery reserve. On-premise and edge inference also keep footage inside your network, which is often the compliance answer.
Your footage stays yours and is used only to train your model. We work under NDA, support on-premise labelling where footage cannot leave the building, and have delivered HIPAA and GDPR compliant systems — VALT across 770+ US organizations, Nucleus under SOC 2 and HIPAA.
You do — the code, the trained weights and the labelled dataset. It is in the contract. Third-party licences and hosting stay in your name and are billed at cost.
Yes, and we start with a free model audit: whether the problem is the data, the pipeline or the thresholds. That distinction usually decides whether the existing work is worth keeping.
The knowledge-base section our engineers hand to clients before the first call.
Face recognition and ANPR, POS-to-video matching, solar cameras on 4G, forensic search.
If you want computer vision engineers inside your own team, we wrote how to do it.
You get back the models worth trying, the cost per stream and an honest answer on whether to build at all. Free, and yours to keep whether you hire us or not.