Introduction

The Zynq-7000 SoC is one of the most practical devices for embedded imaging, because it combines a dual-core Arm Cortex-A9 processor with programmable logic on a single die. The processor runs Linux and the application software, while the fabric implements the image pipeline at deterministic latency. This application note explains how to build that pipeline, from partitioning to memory budgeting and interface design, using the Zynq XC7Z020 as the reference device.

Why Hardware Processing for Vision

Image processing rewards parallelism and determinism. A frame arrives as a stream of pixels, and stages such as demosaic, color correction, filtering and feature extraction can each process many pixels per clock if implemented in hardware. A processor runs these stages sequentially and its timing varies with the operating system load, whereas the fabric executes them at a fixed rate. For machine vision, where repeatability matters as much as raw throughput, that difference is decisive.

Partitioning Software and Hardware

A reliable partition puts the latency-critical stages in the fabric and the flexible stages on the Arm cores. Capture, timing generation, filtering and format conversion belong in hardware; networking, storage, user interface and higher-level decision logic belong on the processor. The boundary is a frame buffer in shared memory, with the fabric writing completed frames and the processor reading them. Keeping that interface narrow and well-defined makes the software easy to maintain.

Streaming Versus Frame-Based

Some stages, such as filtering and edge detection, work naturally on a pixel stream and need only line buffers. Others, such as feature matching or histogram analysis, need a full frame. Put streaming stages first and frame-based stages later, so frames are buffered as late as possible and on-chip memory is used efficiently.

Memory Budgeting

Block RAM is the fastest memory in the device and should hold line buffers and intermediate results. The XC7Z020 provides 4,900 Kb of block RAM and 220 DSP slices, which is enough for a multi-stage pipeline with filtering and simple feature extraction. Full frames live in DDR memory attached to the processor system, and the memory controller bandwidth must cover the frame traffic plus the software's own access. Budget the DDR bandwidth for the worst-case frame rate and resolution.

Bandwidth and Latency

Compute the bandwidth at each stage from the resolution and frame rate, then check it against the memory and interface limits. Where the fabric writes to DDR, use efficient burst transactions to avoid wasting bandwidth. Keep the latency-critical path entirely on-chip so the response time does not depend on DDR access time.

Interface Design

Camera and display interfaces vary widely, and the programmable I/O can be configured for the required standard. Keep high-speed differential pairs short and impedance-controlled, follow the reference stack-up, and route clock and data as matched groups. Confirm the I/O bank voltages and the number of transceivers the design needs before layout, because these constraints are expensive to change later.

Bench Validation

Before the production layout, validate the pipeline on a development board: feed test patterns through each stage, measure the latency end to end, and confirm the frame rate at the target resolution. Injection of synthetic data makes faults repeatable. BeiLuo's FAE team supports capture, filtering and interface questions during design-in and can supply Zynq-7000 samples for bench validation.

Conclusion

Zynq-7000 gives an embedded-vision design both a processor and deterministic hardware in one device. Partition the pipeline carefully, budget the memory against the worst-case frame rate, and validate the interfaces on the bench; do those three things and the design becomes predictable to build and maintain.