Summary
I built and ran every runnable example on an Alveo V80, collected post-route timing for all six hardware designs, and ran each host application.
Results
- Three passed (
01_aximm, 02_chain, 04_freq) and three failed (00_axilite, 06_dcmac, 05_perf).
- Three met timing (
01_aximm, 02_chain, 06_dcmac) and three missed it (00_axilite, 04_freq, 05_perf).
This report covers two findings that look actionable, plus corroborating data for existing issues.
| Example |
Timing |
Host test |
01_aximm |
Met |
Pass |
02_chain |
Met |
Pass |
04_freq |
Missed |
Pass |
00_axilite |
Missed |
Fail (accuracy check, section 2) |
06_dcmac |
Met |
Fail (card management controller died while programming, section 3) |
05_perf |
Missed |
Fail (same card-controller failure, in an earlier run, section 4). Not in the default sweep |
04_freq missed timing and the host test still passed.
06_dcmac met timing and still failed while the card was being programmed.
Headline items:
00_axilite's accuracy check is too strict to pass reliably. The golden model accumulates 1024 floats sequentially on the host and compares against the kernel's accumulator with a 1e-6 relative tolerance. That's tighter than the rounding error inherent in the two different summation orders, so the test fails on a numerically correct result.
- Only three of six designs close timing, and the failures are worth distinguishing from each other: one is confined to debug instrumentation, one is the example asking for a frequency it can't reach, and one is in a static-region clock.
Environment
| Item |
Value |
| Repo commit |
c109e53e566053dd96038619c66224a4d51803a2 (branch off main at 472e5650b65a600d9c5a96580c1ccf680ed69d1b) |
| Vivado / Vitis |
2025.1 |
| Board |
Alveo V80, xcv80-lsva4737-2MHP-e-S |
| AMC |
2.4.0-0.839b4ad.20250821 |
| Packages |
slashkit 1.0.0, vrtd 1.0.0, slash-dkms 1.0.0 |
| OS |
Ubuntu 22.04 |
1. Timing closure across the six hardware examples
Post-route design timing summaries, from each design's impl_1/route_report_timing_summary_0.rpt:
| Example |
Requested clock |
WNS (ns) |
TNS (ns) |
Failing endpoints |
WHS (ns) |
Result |
00_axilite |
200 MHz |
−0.007 |
−1.545 |
359 |
0.000 |
Violated |
01_aximm |
200 MHz |
+0.049 |
0 |
0 |
0.012 |
Met |
02_chain |
200 MHz |
+0.557 |
0 |
0 |
0.008 |
Met |
04_freq |
400 MHz |
−1.290 |
−2968 |
4606 |
0.001 |
Violated |
05_perf |
200 MHz user, 360 MHz static |
−0.244 |
−540.6 |
8079 |
0.000 |
Violated |
06_dcmac |
200 MHz |
+0.074 |
0 |
0 |
0.018 |
Met |
Hold and pulse-width slack are non-negative in all six, with zero failing endpoints, so every failure is setup-only.
The three failures have different characters, which matters for triage:
00_axilite — violation is confined to the debug ILA, not the datapath. All ten worst paths are −0.007 ns and every one terminates inside the ILA that the [debug] section inserts:
Slack (VIOLATED) : -0.007ns
Destination: top_i/slash/axis_ila_debug_0/inst/axis_ila_intf/inst/reg_gap_count_u_reg[3]/R
Destination: top_i/slash/axis_ila_debug_0/inst/axis_ila_intf/inst/reg_trigger_info_complete_reg[35]/R
Destination: top_i/slash/axis_ila_debug_0/inst/axis_ila_intf/inst/reg_trigger_sample_i_reg[376]/R
...
These are reset pins on ILA capture registers. Worth noting because a design whose only timing failure is in instrumentation is a different risk from one failing in the datapath, and because removing the [debug] nets would presumably close it.
04_freq — the example requests more than the design can reach. config.cfg asks for freqhz=400000000 on vadd_0; the longest path is 3.790 ns, so this design tops out near 264 MHz on this device. The runtime already knows and says so at run time, then runs anyway:
Programmed user clock to 263846153 Hz (target 263852242 Hz)
[WARN] void vrt::impl::Device::setFrequency(uint64_t): Setting frequency 300000000, which is higher than max frequency 263852242
The test still reports Test passed. If 400 MHz isn't achievable for this example on V80, it may be worth lowering what the example requests so a clean checkout doesn't produce a design that misses timing by 1.29 ns by default.
05_perf — the bulk of the failure is on a static-region clock. Per-clock breakdown:
| Clock |
Frequency |
WNS (ns) |
TNS (ns) |
Failing endpoints |
top_i/static_region/clk_wizard_0/.../clk_wizard_0_clk_out1_1 |
360 MHz |
−0.244 |
−527.7 |
7988 |
user_clk |
200 MHz |
−0.243 |
−5.4 |
54 |
This looks related to #132, though the correlation with direct HBM isn't exact in my results. Of the four examples using sp=...:HBMn (00_axilite, 01_aximm, 04_freq, 05_perf), 01_aximm closes timing at +0.049 ns; 02_chain and 06_dcmac have no sp= lines at all. 05_perf is also by far the heaviest HBM user here, with 64 channels plus MEM and DDR.
Derived maximum frequencies, if useful: 00_axilite 199.7 MHz, 01_aximm 202.0, 02_chain 225.1, 04_freq 263.9, 05_perf 190.7 on user_clk, 06_dcmac 203.0.
2. 00_axilite accuracy check fails on a numerically correct result
Running the design that was built from an unmodified 00_axilite:
Test failed! (accuracy)
out_r_ctrl: 0x1
Expected: 1536.074707
Got: 1536.072998
Absolute error: 0.001708984375 (effective tolerance 0.001536074677, abs 0.001000000047, rel 9.999999975e-07)
The relative error is about 1.1e-6 against a 1e-6 relative tolerance, so it fails by a hair. The cause looks like the golden model rather than the hardware. In examples/00_axilite/00_axilite.cpp:
uint32_t size = 1024;
...
float goldenModel = 0;
for (uint32_t i = 0; i < size; i++) {
buffer[i] = static_cast<float>(dis(gen));
hostInput[i] = buffer[i];
goldenModel += buffer[i] + 1; // 1024 sequential float adds
}
...
constexpr float kAbsTolerance = 1e-3f;
constexpr float kRelTolerance = 1e-6f;
Accumulating 1024 single-precision values sequentially carries rounding error well above 1e-6 relative; with float epsilon at 1.2e-7, even the random-walk estimate for 1024 terms is around 3.8e-6, and the worst case is far larger. The kernel's accumulator does not sum in the same order as the host loop, so the two results legitimately differ in the last couple of mantissa bits. The test is therefore sensitive to the input random seed, which also explains why it isn't failing for everyone.
I do not think this is data corruption, and I don't think it's the timing violation either: as shown above, every failing path in this design is inside the debug ILA, not the accumulate datapath.
Two straightforward fixes, either of which should do:
double goldenModel = 0; // accumulate the reference in double
or
constexpr float kRelTolerance = 1e-5f; // tolerance appropriate to float accumulation of 1024 terms
Accumulating the reference in double is the more robust of the two, since it keeps the tolerance meaningful as a correctness bound rather than widening it to absorb reference error.
3. 06_dcmac takes down the card management controller
This reproduces #214 and #218 on another host, so I'm recording it as an extra data point rather than a new problem. 06_dcmac fails while programming its service-layer partial PDI:
[INFO] void vrt::impl::Device::programDevice(): Programming PDI via vrtd design writer
.../images/top_i_service_layer_service_layer_dcmac_hw_inst_0_partial.pdi
Exception: Internal error in vrtd daemon or local libvrtd
with the AMC dying at the same moment:
[14:09:27] ami: CRITICAL WARNING: cmd id: 120 timed out(timeout), hot reset is required
[14:09:27] ami 0000:c4:00.0: ERROR : Submitted command timed out
[14:09:27] ami: ERROR : Failed to get the heartbeat msg!
[14:09:27] ami: ERROR : AMC Heartbeat expired event received
[14:09:29] ami: ERROR : Heartbeat fail count above threshold! Raising fatal event...
[14:09:29] ami: ERROR : AMC Heartbeat fatal event received, stopping GCQ...
Recovery needed sudo ami_tool reload plus a vrtd restart. Note that v80-smi list still reports the board fully healthy while the AMC is down, which makes this harder to diagnose than it needs to be:
Board 0000:c4:00 OK (PF0: OK) (PF1: OK) (PF2: OK) (VRTD: OK) Shell: service
If v80-smi list could surface AMC heartbeat state, the failure would be self-explanatory instead of surfacing as a generic Internal error in vrtd daemon or local libvrtd.
4. 05_perf fails the same way
Earlier the same day, 05_perf produced an identical AMC failure while programming its PDI:
[13:29:33] ami 0000:c4:00.0: ERROR : Logging thread error - invalid PCI data read!
[13:29:34] ami: CRITICAL WARNING: cmd id: 12 timed out(timeout), hot reset is required
[13:29:34] ami: ERROR : AMC Heartbeat expired event received
[13:29:37] ami: ERROR : AMC Heartbeat fatal event received, stopping GCQ...
So two of the six designs, 05_perf and 06_dcmac, reliably require a card reset after being loaded on this system.
Reproduction tooling
Two scripts I used, in case they're useful to others or worth upstreaming in some form.
examples/build_example.sh — build one example end to end
Takes an example name, optionally cleans, configures, builds the host app, compiles the HLS kernels serially, then links the vbin. Two notes on the less obvious parts:
- It runs the HLS target with
--parallel 1. Two concurrent v++ --mode hls jobs contend on the shared Vitis part database and fail with HLS 200-1608 Failed to open platform database.
- The
platform_db_readable probe and lock shim exist because the machine I was on had a broken NFS lock manager, where every fcntl lock returned ENOLCK and SQLite therefore reported the read-only platform database as locked. That's a host problem rather than a SLASH one, and the block is skipped entirely on a healthy machine; it's included only so the script reads coherently.
SLASH_LINK_JOBS needs the repo's cmake/SlashTools.cmake to take effect, since the installed SlashTools package predates the JOBS parameter and silently leaves slashkit at its default of 8 Vivado jobs.
#!/bin/bash
# -e: exit on command failure
# -u: error on unset variables
# -o pipefail: fail a pipeline if any command in it fails
set -euo pipefail
# cd to this script's directory so cmake runs here even if invoked from elsewhere
cd "$(dirname "$0")"
usage() {
echo "Usage: $0 <example> [-c]" >&2
echo " -c remove <example>/build/ before building" >&2
echo " Example: $0 00_axilite -c" >&2
echo " Build options:" >&2
for d in */; do
if [[ -f "${d}CMakeLists.txt" ]]; then
echo " ${d%/}" >&2
fi
done
}
CLEAN=0
EXAMPLE=""
for arg in "$@"; do
case "$arg" in
-c)
CLEAN=1
;;
-h|--help)
usage
exit 0
;;
-*)
echo "Error: unknown option: $arg" >&2
usage
exit 1
;;
*)
if [[ -n "$EXAMPLE" ]]; then
echo "Error: unexpected argument: $arg" >&2
usage
exit 1
fi
EXAMPLE="$arg"
;;
esac
done
if [[ -z "$EXAMPLE" ]]; then
usage
exit 1
fi
if [[ ! -d "$EXAMPLE" ]]; then
echo "Error: example directory not found: $EXAMPLE" >&2
exit 1
fi
if [[ ! -f "$EXAMPLE/CMakeLists.txt" ]]; then
echo "Error: not a cmake example: $EXAMPLE" >&2
usage
exit 1
fi
# Use the add_vbin hw target from CMakeLists.txt (dir name may differ, e.g.
# 00_axilite_200mhz still builds axilite_hw). Examples without one, such as
# 03_multiple_boards, are host-only: they run a prebuilt vrtbin, so there is
# nothing to synthesise and only the host application gets built.
HW_TARGET="$(sed -n 's/.*add_vbin(TARGET "\([^"]*_hw\)".*/\1/p' "$EXAMPLE/CMakeLists.txt" | head -n 1)"
cd "$EXAMPLE"
if [[ "$CLEAN" -eq 1 ]]; then
echo "-----------------------------------------------------"
echo "Removing build/"
rm -rf build/
fi
echo "-----------------------------------------------------"
echo "Building $EXAMPLE"
echo "Running cmake"
cmake -B build -S . -G Ninja
echo "-----------------------------------------------------"
echo ">>> Building with Ninja for $EXAMPLE >>>"
cmake --build build
echo " Building with Ninja complete for $EXAMPLE"
if [[ -z "$HW_TARGET" ]]; then
echo "-----------------------------------------------------"
echo "$EXAMPLE is host-only: no HLS kernels and no vbin to link."
echo "It runs a prebuilt vrtbin, so build the hardware in the example"
echo "that provides it and make sure the vrtbin is on the run path."
exit 0
fi
echo "-----------------------------------------------------"
echo ">>> Compiling HLS kernels for $EXAMPLE >>>"
# Vitis HLS uses a shared part database; two v++ --mode hls jobs at once
# fail with "database is locked" (HLS 200-1608).
cmake --build build --target hls --parallel 1
echo " Compiling HLS kernels complete for $EXAMPLE"
echo "-----------------------------------------------------"
echo ">>> Linking into a hardware verilog binary for $EXAMPLE >>>"
echo " Tail build log: "
echo " tail -f build/${HW_TARGET}.vbin.prj/logs/slash_project_build.log"
echo " Building..."
cmake --build build --target "$HW_TARGET" # link into a hardware verilog binary
echo " Linking into a hardware verilog binary complete for $EXAMPLE"
#echo "-----------------------------------------------------"
#echo " "
#echo " Finding hardware verilog binaries for $EXAMPLE"
#find . -name "*.vbin"
#echo " "
#echo " Listing hardware verilog binaries for $EXAMPLE"
#v80-smi list
examples/run_all.sh — run every example's host application
Works out which static shell each design was linked against by reading static_shell versus static_shell_compute out of its slash_project_build.log, groups the tests accordingly, and starts with whichever group matches the shell already loaded so a full sweep needs one v80-smi reset --shell-type rather than two. It aborts the remaining tests when a failure coincides with an AMC heartbeat fatal event in dmesg, because after that every later test fails identically. 05_perf is opt-in via --with-perf for the reason in section 4.
It also refuses to start when v80-smi list returns nothing. That happens once PF0 has fallen off the PCIe bus, after which v80-smi reset hangs rather than failing, so the script would otherwise sit there until interrupted.
#!/bin/bash
# -e: exit on command failure
# -u: error on unset variables
# -o pipefail: fail a pipeline if any command in it fails
set -euo pipefail
# cd to this script's directory so the relative example paths below resolve even
# when invoked from elsewhere
cd "$(dirname "$0")"
DEVICE="c4:00"
INCLUDE_PERF=0
PERF_KERNELS=1
LOG_DIR="${TMPDIR:-/tmp}/slash-run-all-$UID"
usage() {
echo "Usage: $0 [-d <bdf>] [--with-perf [n]] [<example> ...]" >&2
echo " -d <bdf> board address (default: $DEVICE)" >&2
echo " --with-perf [n] also run 05_perf, with n kernels (default: $PERF_KERNELS)" >&2
echo " <example> ... run only these, instead of every runnable example" >&2
echo " Runs each example's host application against its linked vbin," >&2
echo " switching the board shell when the next test needs the other one." >&2
echo "" >&2
echo " Examples:" >&2
echo " $0" >&2
echo " $0 00_axilite" >&2
echo " $0 00_axilite --with-perf 1" >&2
echo " $0 00_axilite --with-perf 1 01_example" >&2
echo " $0 00_axilite --with-perf 1 01_example 02_udp" >&2
echo " $0 00_axilite --with-perf 1 01_example 02_udp 03_multiple_boards" >&2
echo " $0 00_axilite --with-perf 1 01_example 02_udp 03_multiple_boards 04_udp" >&2
echo " $0 00_axilite --with-perf 1 01_example 02_udp 03_multiple_boards 04_udp 05_perf" >&2
echo " $0 00_axilite --with-perf 1 01_example 02_udp 03_multiple_boards 04_udp 05_perf 06_udp" >&2
}
ONLY=()
while [[ $# -gt 0 ]]; do
case "$1" in
-d)
[[ $# -ge 2 ]] || { echo "Error: -d needs a board address" >&2; exit 1; }
DEVICE="$2"
shift 2
;;
--with-perf)
INCLUDE_PERF=1
if [[ ${2:-} =~ ^[0-9]+$ ]]; then
PERF_KERNELS="$2"
shift
fi
shift
;;
-h|--help)
usage
exit 0
;;
-*)
echo "Error: unknown option: $1" >&2
usage
exit 1
;;
*)
ONLY+=("$1")
shift
;;
esac
done
command -v v80-smi >/dev/null || { echo "Error: v80-smi not on PATH; source setup_env" >&2; exit 1; }
# An example is runnable when it has both a host executable and a linked vbin.
# 03_multiple_boards and 07_udp never qualify: neither produces a vbin here.
vbin_of() {
ls "$1"/build/*_hw.vbin 2>/dev/null | head -n 1
}
# Usually build/<example>, but the executable takes its name from the cmake
# project, which does not always match the directory (03_multiple_boards builds
# 03_example), so fall back to the only executable cmake left in build/.
exe_of() {
if [[ -x "$1/build/$1" ]]; then
echo "$1/build/$1"
return
fi
find "$1/build" -maxdepth 1 -type f -executable -not -name '*.cmake' \
-not -name '*.ninja' 2>/dev/null | head -n 1
}
# The static shell a design was linked against decides which shell the board has
# to boot, and it can hold only one at a time, so the tests are grouped by it.
# The build log records the checkpoint slashkit passed to Vivado.
shell_of() {
local log
log="$(find "$1/build" -name slash_project_build.log 2>/dev/null | head -n 1)"
if [[ -z "$log" ]]; then
echo unknown
elif grep -q 'static_shell_compute/static_shell_slash.dcp' "$log"; then
echo compute
else
echo service
fi
}
board_shell() {
v80-smi list 2>/dev/null | sed -n 's/.*Shell: *\([a-z]*\).*/\1/p' | head -n 1
}
# A dead card management controller makes every later test fail identically, so
# stop rather than hammer it. The driver asks for a hot reset when this happens.
amc_is_dead() {
dmesg -T 2>/dev/null | grep -i 'ami' | tail -n 20 \
| grep -q 'AMC Heartbeat fatal event received'
}
CANDIDATES=()
if [[ ${#ONLY[@]} -gt 0 ]]; then
for d in "${ONLY[@]}"; do
if [[ -z "$(exe_of "$d")" ]]; then
echo "Error: no host executable under $d/build; run ./build.sh $d" >&2
exit 1
fi
if [[ -z "$(vbin_of "$d")" ]]; then
echo "Error: no linked vbin under $d/build; run ./build.sh $d" >&2
exit 1
fi
CANDIDATES+=("$d")
done
else
for d in */; do
d="${d%/}"
[[ -n "$(exe_of "$d")" && -n "$(vbin_of "$d")" ]] || continue
if [[ "$d" == "05_perf" && "$INCLUDE_PERF" -eq 0 ]]; then
continue
fi
CANDIDATES+=("$d")
done
fi
if [[ ${#CANDIDATES[@]} -eq 0 ]]; then
echo "Error: nothing runnable found; build an example first" >&2
exit 1
fi
# Start with whichever group matches the shell already loaded, so a full run
# needs one switch instead of two.
CURRENT="$(board_shell)"
SERVICE=()
COMPUTE=()
for ex in "${CANDIDATES[@]}"; do
case "$(shell_of "$ex")" in
compute) COMPUTE+=("$ex") ;;
service) SERVICE+=("$ex") ;;
*) echo "Skipping $ex: cannot tell which shell it was linked against" >&2 ;;
esac
done
ORDER=()
if [[ "$CURRENT" == "compute" ]]; then
ORDER=("${COMPUTE[@]:-}" "${SERVICE[@]:-}")
else
ORDER=("${SERVICE[@]:-}" "${COMPUTE[@]:-}")
fi
# No point starting, or trying to switch shells, when the board cannot be read.
# v80-smi prints nothing at all once PF0 has fallen off the PCIe bus, and the
# reset it would otherwise attempt hangs instead of failing.
if [[ -z "$CURRENT" ]]; then
echo "Error: v80-smi reports no board; PF0 may have dropped off the bus" >&2
echo " lspci -s $DEVICE # expect functions .0, .1 and .2" >&2
echo " dmesg -T | grep -i ami | tail" >&2
exit 1
fi
mkdir -p "$LOG_DIR"
RESULTS=()
ABORTED=""
for ex in "${ORDER[@]}"; do
[[ -n "$ex" ]] || continue
WANT="$(shell_of "$ex")"
CURRENT="$(board_shell)"
if [[ "$WANT" != "$CURRENT" ]]; then
echo "-----------------------------------------------------"
echo "Booting the $WANT shell for $ex (board is on ${CURRENT:-unknown})"
if ! v80-smi reset -d "$DEVICE" --shell-type "$WANT"; then
ABORTED="v80-smi reset to the $WANT shell failed"
break
fi
CURRENT="$(board_shell)"
if [[ "$CURRENT" != "$WANT" ]]; then
ABORTED="board reports the $CURRENT shell after asking for $WANT"
break
fi
fi
EXE="$(exe_of "$ex")"
VBIN="$(vbin_of "$ex")"
ARGS=("$DEVICE" "$VBIN")
if [[ "$ex" == "05_perf" ]]; then
ARGS+=("$PERF_KERNELS")
fi
LOG="$LOG_DIR/$ex.log"
echo "-----------------------------------------------------"
echo ">>> $ex on the $WANT shell >>>"
echo " $EXE ${ARGS[*]}"
if "./$EXE" "${ARGS[@]}" >"$LOG" 2>&1; then
RESULTS+=("PASS $ex")
echo " PASS"
else
RESULTS+=("FAIL $ex")
echo " FAIL, see $LOG"
tail -n 3 "$LOG" | sed 's/^/ /'
# Programming failures leave the card unusable for every later test, so
# there is nothing to learn from carrying on.
if grep -q 'Internal error in vrtd daemon' "$LOG" && amc_is_dead; then
ABORTED="the card's AMC stopped responding; it needs a hot reset"
break
fi
fi
done
echo "-----------------------------------------------------"
echo "Results"
for r in "${RESULTS[@]:-}"; do
[[ -n "$r" ]] && echo " $r"
done
echo " logs: $LOG_DIR"
if [[ -n "$ABORTED" ]]; then
echo "-----------------------------------------------------" >&2
echo "Stopped early: $ABORTED" >&2
echo " sudo ami_tool reload -t sbr -d $DEVICE # then pci, then driver" >&2
echo " sudo systemctl restart vrtd && v80-smi list" >&2
exit 1
fi
for r in "${RESULTS[@]:-}"; do
[[ "$r" == FAIL* ]] && exit 1
done
exit 0
Related issues
Summary
I built and ran every runnable example on an Alveo V80, collected post-route timing for all six hardware designs, and ran each host application.
Results
01_aximm,02_chain,04_freq) and three failed (00_axilite,06_dcmac,05_perf).01_aximm,02_chain,06_dcmac) and three missed it (00_axilite,04_freq,05_perf).This report covers two findings that look actionable, plus corroborating data for existing issues.
01_aximm02_chain04_freq00_axilite06_dcmac05_perf04_freqmissed timing and the host test still passed.06_dcmacmet timing and still failed while the card was being programmed.Headline items:
00_axilite's accuracy check is too strict to pass reliably. The golden model accumulates 1024floats sequentially on the host and compares against the kernel's accumulator with a 1e-6 relative tolerance. That's tighter than the rounding error inherent in the two different summation orders, so the test fails on a numerically correct result.Environment
c109e53e566053dd96038619c66224a4d51803a2(branch offmainat472e5650b65a600d9c5a96580c1ccf680ed69d1b)xcv80-lsva4737-2MHP-e-S2.4.0-0.839b4ad.20250821slashkit 1.0.0,vrtd 1.0.0,slash-dkms 1.0.01. Timing closure across the six hardware examples
Post-route design timing summaries, from each design's
impl_1/route_report_timing_summary_0.rpt:00_axilite01_aximm02_chain04_freq05_perf06_dcmacHold and pulse-width slack are non-negative in all six, with zero failing endpoints, so every failure is setup-only.
The three failures have different characters, which matters for triage:
00_axilite— violation is confined to the debug ILA, not the datapath. All ten worst paths are −0.007 ns and every one terminates inside the ILA that the[debug]section inserts:These are reset pins on ILA capture registers. Worth noting because a design whose only timing failure is in instrumentation is a different risk from one failing in the datapath, and because removing the
[debug]nets would presumably close it.04_freq— the example requests more than the design can reach.config.cfgasks forfreqhz=400000000onvadd_0; the longest path is 3.790 ns, so this design tops out near 264 MHz on this device. The runtime already knows and says so at run time, then runs anyway:The test still reports
Test passed. If 400 MHz isn't achievable for this example on V80, it may be worth lowering what the example requests so a clean checkout doesn't produce a design that misses timing by 1.29 ns by default.05_perf— the bulk of the failure is on a static-region clock. Per-clock breakdown:top_i/static_region/clk_wizard_0/.../clk_wizard_0_clk_out1_1user_clkThis looks related to #132, though the correlation with direct HBM isn't exact in my results. Of the four examples using
sp=...:HBMn(00_axilite,01_aximm,04_freq,05_perf),01_aximmcloses timing at +0.049 ns;02_chainand06_dcmachave nosp=lines at all.05_perfis also by far the heaviest HBM user here, with 64 channels plusMEMandDDR.Derived maximum frequencies, if useful:
00_axilite199.7 MHz,01_aximm202.0,02_chain225.1,04_freq263.9,05_perf190.7 onuser_clk,06_dcmac203.0.2.
00_axiliteaccuracy check fails on a numerically correct resultRunning the design that was built from an unmodified
00_axilite:The relative error is about 1.1e-6 against a 1e-6 relative tolerance, so it fails by a hair. The cause looks like the golden model rather than the hardware. In
examples/00_axilite/00_axilite.cpp:Accumulating 1024 single-precision values sequentially carries rounding error well above 1e-6 relative; with
floatepsilon at 1.2e-7, even the random-walk estimate for 1024 terms is around 3.8e-6, and the worst case is far larger. The kernel's accumulator does not sum in the same order as the host loop, so the two results legitimately differ in the last couple of mantissa bits. The test is therefore sensitive to the input random seed, which also explains why it isn't failing for everyone.I do not think this is data corruption, and I don't think it's the timing violation either: as shown above, every failing path in this design is inside the debug ILA, not the accumulate datapath.
Two straightforward fixes, either of which should do:
or
Accumulating the reference in
doubleis the more robust of the two, since it keeps the tolerance meaningful as a correctness bound rather than widening it to absorb reference error.3.
06_dcmactakes down the card management controllerThis reproduces #214 and #218 on another host, so I'm recording it as an extra data point rather than a new problem.
06_dcmacfails while programming its service-layer partial PDI:with the AMC dying at the same moment:
Recovery needed
sudo ami_tool reloadplus avrtdrestart. Note thatv80-smi liststill reports the board fully healthy while the AMC is down, which makes this harder to diagnose than it needs to be:If
v80-smi listcould surface AMC heartbeat state, the failure would be self-explanatory instead of surfacing as a genericInternal error in vrtd daemon or local libvrtd.4.
05_perffails the same wayEarlier the same day,
05_perfproduced an identical AMC failure while programming its PDI:So two of the six designs,
05_perfand06_dcmac, reliably require a card reset after being loaded on this system.Reproduction tooling
Two scripts I used, in case they're useful to others or worth upstreaming in some form.
examples/build_example.sh— build one example end to endTakes an example name, optionally cleans, configures, builds the host app, compiles the HLS kernels serially, then links the vbin. Two notes on the less obvious parts:
--parallel 1. Two concurrentv++ --mode hlsjobs contend on the shared Vitis part database and fail withHLS 200-1608 Failed to open platform database.platform_db_readableprobe and lock shim exist because the machine I was on had a broken NFS lock manager, where everyfcntllock returnedENOLCKand SQLite therefore reported the read-only platform database as locked. That's a host problem rather than a SLASH one, and the block is skipped entirely on a healthy machine; it's included only so the script reads coherently.SLASH_LINK_JOBSneeds the repo'scmake/SlashTools.cmaketo take effect, since the installedSlashToolspackage predates theJOBSparameter and silently leaves slashkit at its default of 8 Vivado jobs.examples/run_all.sh— run every example's host applicationWorks out which static shell each design was linked against by reading
static_shellversusstatic_shell_computeout of itsslash_project_build.log, groups the tests accordingly, and starts with whichever group matches the shell already loaded so a full sweep needs onev80-smi reset --shell-typerather than two. It aborts the remaining tests when a failure coincides with an AMC heartbeat fatal event indmesg, because after that every later test fails identically.05_perfis opt-in via--with-perffor the reason in section 4.It also refuses to start when
v80-smi listreturns nothing. That happens once PF0 has fallen off the PCIe bus, after whichv80-smi resethangs rather than failing, so the script would otherwise sit there until interrupted.Related issues
sp=HBMdesigns, including05_perf. Section 1 adds per-clock data; note that01_aximmusessp=...:HBM0and does close timing here.06_dcmacstalls while programming the service-layer PDI and leaves the board requiring reset #214 and VRTD internal error on DCMAC example #218 —06_dcmacservice-layer PDI failure. Section 3 is another occurrence.00_axiliteproblem (needing two runs). Section 2 is unrelated to it.