This repository has been archived on 2026-02-22 . You can view files and clone it. You cannot open issues or pull requests or push a commit.
041968f38a830def13c0aad2070297817340cda2
tkDNN
tkDNN is a Deep Neural Network library built with cuDNN primitives specifically thought to work on NVIDIA TK1(and all successive) board.
The main scope is to do high performance inference on already trained models.
this branch actually work on every NVIDIA GPU that support the dependencies:
- CUDA 9
- CUDNN 7.105
- TENSORRT 4.02
Workflow
The recommended workflow follow these step:
- Build and train a model in Keras (on any PC)
- Export weights and bias
- Define the model on tkDNN
- Do inference (on TK1)
Compile the library
Build with cmake
mkdir build
cd build
cmake ..
# use -DTEST_DATA=False to skip dataset download
make
during the cmake configuration it will be dowloaded the weights needed for running the tests
Test
Assumiung you have correctly builded the library these are the test ready to exec:
- test_simple: a simple convolutional and dense network (CUDNN only)
- test_mnist: the famous mnist netwok (CUDNN and TENSORRT)
- test_mnistRT: the mnist network hardcoded in using tensorRT apis (TENSORRT only)
- test_yolo: YOLO detection network (CUDNN and TENSORRT)
- test_yolo_tiny: smaller version of YOLO (CUDNN and TENSRRT)
- test_yolo3_berkeley: our yolo3 version trained with BDD100K dateset
yolo3 berkeley demo detection
For the live detection you need to precompile the tensorRT file by luncing the desidered network test, this is the recommended process:
export TKDNN_MODE=FP16 # set the half floating point optimization
rm yolo3_berkeley.rt # be sure to delete(or move) old tensorRT files
./test_yolo3_berkeley # run the yolo test (is slow)
# with f16 inference the result will be a bit incorrect
this will genereate a yolo3_berkeley.rt file that can be used for live detection:
./yolo3_demo # launch detection on a demo video
./yolo3_demo yolo3_berkeley.rt /dev/video0 # launch detection on device 0
Releases
4
version 0.6
Latest
Languages
C++
90.9%
Cuda
4%
Python
2.3%
Shell
1.1%
CMake
1.1%
Other
0.6%