Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
53 changes: 47 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,11 +1,52 @@
**University of Pennsylvania, CIS 5650: GPU Programming and Architecture,
Project 1 - Flocking**

* (TODO) YOUR NAME HERE
* (TODO) [LinkedIn](), [personal website](), [twitter](), etc.
* Tested on: (TODO) Windows 22, i7-2222 @ 2.22GHz 22GB, GTX 222 222MB (Moore 2222 Lab)
* Jiahang Mao
* [LinkedIn](https://www.linkedin.com/in/jay-jiahang-m-b05608192/)
* Tested on: Windows 11, i5-13600kf @ 5.0GHz 64GB, RTX 4090 24GB, Personal Computer

### (TODO: Your README)
## Visual Results
5000 boids
![Boids Simulation](images/out.gif)

Include screenshots, analysis, etc. (Remember, this is public, so don't put
anything here that you don't want to share with the world.)


## Performance Results

Performance testing config
* Avg FPS measured between 2 seconds and 7 seconds after program start, calculated as total_fps / total_duration
* run in Release model

#### Visualization OFF

| Method | Number of Boids | Block Size | Avg FPS |
|:-------- |:---------------:|:----------:|:-------:|
| Naive (baseline) | 5000 | 128 | 878
| Naive | 5000 | 1024 | 603
| Naive | 50000 | 128 | 135
| Naive | 50000 | 1024 | 86
| Uniform | 5000 | 128 | 1582
| Uniform | 5000 | 1024 | 1617
| Uniform | 50000 | 128 | 1279
| Uniform | 50000 | 1024 | 1213
| Coherent | 5000 | 128 | 1571
| Coherent | 5000 | 1024 | 1555
| Coherent | 50000 | 128 | 1027
| Coherent | 50000 | 1024 | 1030

#### Visualization ON

| Method | Number of Boids | Block Size | Avg FPS |
|:-------- |:---------------:|:----------:|:-------:|
| Naive (baseline) | 5000 | 128 | 715
| Naive | 5000 | 1024 | 526
| Naive | 50000 | 128 | 130
| Naive | 50000 | 1024 | 84
| Uniform | 5000 | 128 | 1226
| Uniform | 5000 | 1024 | 1133
| Uniform | 50000 | 128 | 961
| Uniform | 50000 | 1024 | 880
| Coherent | 5000 | 128 | 1042
| Coherent | 5000 | 1024 | 1070
| Coherent | 50000 | 128 | 764
| Coherent | 50000 | 1024 | 773
Binary file added images/out.gif
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
18 changes: 18 additions & 0 deletions p3_writeup.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@

## Part3 write-up
* For each implementation, how does changing the number of boids affect performance? Why do you think this is?

As the number of boids increase, performance drop. Because more boids would require more frequent access to memory for accessing related position and velocity informatoin. Also once the number of boids exceed the maximum number of threads supported by hardware, we need more gpu cycles to run all boids.
* For each implementation, how does changing the block count and block size affect performance? Why do you think this is?

According to my experiments. Increasing the block count causes worse performance in naive implementation ( 25% less fps). While for uniform and coherent grids it barely has any impact on performance. <br>
For naive implementation, i think it was because more branches occured inside one block, causing wasted cycles.
For coherent and uniform grid, the branching is minizied since the number of neighbours to check decreased on each thread.

* For the coherent uniform grid: did you experience any performance improvements with the more coherent uniform grid? Was this the outcome you expected? Why or why not?

I saw slightly worse performance with coherent uniform grids. Maybe my implementation is faulty. I suspect the additional sorting required more time, and the memory coherency isn't utilized to local SM cache partitions

* Did changing cell width and checking 27 vs 8 neighboring cells affect performance? Why or why not? Be careful: it is insufficient (and possibly incorrect) to say that 27-cell is slower simply because there are more cells to check!

I observed checking 27 performs better than 8 about 2-3% with both uniform and coherent grids. I suppose because the total volume to check is smaller with 27, since 27 means checking 3^3 unit volumes while there is a chance that 8 end up checking (2*2)^3 unit volumes. But the difference is marginal.
Loading