summaryrefslogtreecommitdiff
path: root/report.org
blob: 044e888ab8acd6c42c6656cf6d6741beb6b2363c (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
#+title: Diabetic Retinopathy Classification -- Final Report
#+subtitle: EEL4759: Digital Image Processing
#+author: Vineet Kumar
#+options: toc:nil
#+cite_export: biblatex
#+bibliography: references.bib
#+latex_class_options: [12pt,letterpage]
#+latex_header: \usepackage[margin=1in]{geometry}
#+latex_header: \usepackage{booktabs}
#+latex_header: \usepackage{float}
#+latex_header: \usepackage{graphicx}

* Abstract

This report covers automated severity grading of diabetic retinopathy
(DR) from retinal images using deep convolutional neural
networks. Three pretrained architectures (ResNet-50, EfficientNet-B0,
and ViT-B/16) were fine-tuned on the Kaggle "Diabetic Retinopathy
Resized Arranged" dataset and evaluated on a five-class severity
classification task (Healthy through Proliferative DR). Each model was
trained with and without CLAHE (Contrast Limited Adaptive Histogram
Equalization) preprocessing to assess its effect on classification
performance. Class imbalance was addressed via sqrt-inverse-frequency
weighted sampling and Focal Loss. At 384\times{}384 pixel resolution,
ResNet-50 and EfficientNet-B0 achieved 0.80 overall accuracy and
weighted F1-scores of 0.77 and 0.78 respectively. ViT-B/16 was
constrained to 224\times{}224 and achieved 0.69 accuracy. CLAHE
improved detection of minority classes (particularly Mild NPDR) at a
small cost to overall accuracy. Mild NPDR remained the hardest class
to classify across all experiments, with F1-scores no higher than 0.17
due to its visual similarity to healthy retinal images.

* Introduction

** Diabetic Retinopathy

Diabetic retinopathy is a progressive eye disease caused by diabetes
that is one of the leading causes of preventable vision loss
worldwide. It develops when high blood sugar damages the retinal blood
vessels, causing leakage, abnormal blood vessel growth, and eventually
retinal detachment if untreated. The disease is classified into five
severity grades: no retinopathy (Healthy), Mild Non-Proliferative DR
(NPDR), Moderate NPDR, Severe NPDR, and Proliferative DR. Early
detection through systematic retinal screening can prevent up to 90%
of severe vision loss cases. Manual grading by trained specialists is
accurate but expensive and slow, making automated image classification
useful for large-scale screening programs.

** Convolutional Neural Networks for Retinal Image Classification

Convolutional neural networks (CNNs) have become the dominant approach
for medical image classification tasks. Rather than hand-crafting
features, CNNs learn hierarchical representations (low-level edges and
textures in early layers, progressively more abstract features in
deeper layers) directly from training data. For retinal imaging,
pretrained ImageNet models (transfer learning) are effective: the
low-level filters learned on natural images transfer well to retinal
images, and fine-tuning on DR data allows the network to specialize to
lesion-specific features such as microaneurysms, hard exudates, and
hemorrhages that distinguish severity grades.  This project compares
three architectures spanning different design philosophies: ResNet-50
(residual connections), EfficientNet-B0 (compound scaling), and
ViT-B/16 (pure self-attention on image patches).

** Existing Model Performance

The Kaggle Diabetic Retinopathy detection competitions have
established rough baselines for this type of task. On five-class
grading datasets similar to the one used here, top-performing single
models typically achieve quadratic weighted kappa scores in the
0.82--0.86 range using heavily tuned ensembles, test-time
augmentation, and competition-grade preprocessing. More comparable
single-model baselines reported on Kaggle notebooks for the "Diabetic
Retinopathy Resized Arranged" dataset (224\times{}224) report
accuracies of roughly 0.60--0.73 and kappa values around 0.50--0.58
for standard fine-tuned CNNs, which is consistent with the results
obtained in this project's 224px training run.

* Methodology

** Dataset and Class Distribution

The dataset used is "Diabetic Retinopathy Resized Arranged"
[cite:@amanneo2022dr], containing 35,126 retinal images organized into
five class-labeled folders. The images are JPEG files at approximately
1024\times{}768 pixel resolution. The dataset was split into 70%
training, 15% validation, and 15% test sets using stratified sampling
to preserve the class distribution across all splits (scikit-learn's
=train_test_split= with =stratify=, seed 42).

#+begin_export latex
\begin{table}[H]
\centering
\caption{Dataset class distribution (test split, 5,269 images).}
\begin{tabular}{clrr}
\toprule
Class & Grade & Test Count & Test \% \\
\midrule
0 & Healthy          & 3,872 & 73.5\% \\
1 & Mild NPDR        &   366 &  6.9\% \\
2 & Moderate NPDR    &   794 & 15.1\% \\
3 & Severe NPDR      &   131 &  2.5\% \\
4 & Proliferative DR &   106 &  2.0\% \\
\bottomrule
\end{tabular}
\end{table}
#+end_export

The dataset is severely imbalanced: the Healthy class accounts for
nearly 74% of all samples, while Severe NPDR and Proliferative DR
together comprise only 4.5%.

** Models

Three ImageNet-pretrained architectures were fine-tuned for 5-class DR
severity classification. In all cases the original classification head
was replaced with a linear layer mapping to 5 outputs, and all
backbone weights were unfrozen for full fine-tuning.

- ResNet-50 [cite:@he2016deep]: Deep residual network with skip
  connections. The final fully connected layer (2048 \to 5) was
  replaced. Input: 384\times{}384.

- EfficientNet-B0 [cite:@tan2019efficientnet]: Compound-scaled network
  optimizing width, depth, and resolution simultaneously. The
  classifier head (=Dropout(0.2)= + Linear 1280 \to 5) was replaced.
  Input: 384\times{}384.

- ViT-B/16 [cite:@dosovitskiy2021image]: Vision Transformer that
  splits the image into 16\times{}16 pixel patches and processes them
  with multi-head self-attention. The projection head (768 \to 5) was
  replaced. Input: 224\times{}224 (PyTorch's ViT-B/16 implementation
  only supports 224\times{}224 input; higher resolutions are not
  supported without custom modifications to the model).

** Preprocessing and Augmentation

Images were resized from their native ~1024\times{}768 resolution to
the model's target input resolution (384\times{}384 for ResNet-50 and
EfficientNet-B0; 224\times{}224 for ViT-B/16) as the first
preprocessing step.

CLAHE (Contrast Limited Adaptive Histogram Equalization)
[cite:@zuiderveld1994clahe] was applied as an optional subsequent
preprocessing step. The transform converts the image from RGB to LAB
color space, applies =cv2.createCLAHE= (clip limit 2.0, tile grid
8\times{}8) to the L (lightness) channel only, and converts back to
RGB. Operating on the lightness channel alone
enhances local contrast in retinal structures such as vessels and
lesions without altering hue or saturation.

The training augmentation pipeline (applied after CLAHE if enabled):

1. =RandomResizedCrop= to target resolution, scale 0.8--1.0
2. =RandomHorizontalFlip= (p=0.5), =RandomVerticalFlip= (p=0.5)
3. =RandomRotation=(\pm{}30\textdegree)
4. =ColorJitter= (brightness 0.3, contrast 0.3, saturation 0.2, hue
   0.02)
5. =GaussianBlur= (kernel 3, \sigma{} 0.1--1.0)
6. Normalize to ImageNet mean/std

Validation and test images were resized to the target resolution
deterministically (no random crop) and normalized identically.

** Class Imbalance Handling

Two complementary mechanisms addressed the severe class imbalance:

- WeightedRandomSampler: Each training sample was assigned a weight
  proportional to $1/\sqrt{n_c}$ where $n_c$ is the count of its
  class. The square-root weighting provides a moderate upsampling of
  minority classes without completely drowning the majority-class
  signal (which full inverse-frequency weighting was found to do).

- Focal Loss [cite:@lin2017focal]: The loss function
  $\mathrm{FL}(p_t) = -(1-p_t)^\gamma \log(p_t)$ with $\gamma=2$
  down-weights easy, well-classified examples and focuses gradient
  updates on hard misclassifications, which helps with minority classes
  where the model initially predicts low confidence.

** Training Configuration

#+begin_export latex
\begin{table}[H]
\centering
\caption{Hyperparameters used for all experiments.}
\begin{tabular}{ll}
\toprule
Hyperparameter & Value \\
\midrule
Optimizer              & AdamW \\
Head learning rate     & 1e-4 \\
Backbone learning rate & 1e-5 ($10\times$ lower) \\
Weight decay           & 1e-4 \\
Batch size             & 256 \\
Max epochs             & 50 \\
Early stopping         & patience 10, monitor val macro F1 \\
LR warmup              & 3 epochs linear ($0.1\times \to 1\times$) \\
LR schedule            & Cosine annealing after warmup \\
Gradient clipping      & max norm 1.0 \\
Mixed precision        & \texttt{torch.amp.autocast} + GradScaler \\
Hardware               & $2\times$ AMD Radeon RX 7900 XTX (ROCm) \\
Random seed            & 42 \\
\bottomrule
\end{tabular}
\end{table}
#+end_export

A discriminative learning rate was used: the pretrained backbone was
trained at 1e-5 while the newly initialized classification head was
trained at the full 1e-4. This prevents the pretrained features from
being overwritten too quickly in early epochs.

** Evaluation Metrics

Each trained model was evaluated on the held-out test set using:

- Accuracy: fraction of correctly classified samples
- Weighted F1: F1 averaged over classes weighted by support; reflects
  overall performance on the imbalanced distribution
- Macro F1: F1 averaged equally over all 5 classes; better reflects
  performance on minority classes
- Quadratic-weighted Cohen's Kappa: measures inter-rater agreement
  beyond chance; the quadratic weighting penalizes predictions further
  from the true label more heavily, which is appropriate for the
  ordinal DR severity scale

* Results

** Summary of Test-Set Performance

#+begin_export latex
\begin{table}[H]
\centering
\caption{Test-set metrics for all six experiments at $384\times384$ (ResNet-50, EfficientNet-B0) and $224\times224$ (ViT-B/16).}
\begin{tabular}{llccc}
\toprule
Model & CLAHE & Accuracy & Weighted F1 & Macro F1 \\
\midrule
ResNet-50       & No  & 0.80 & 0.77 & 0.51 \\
ResNet-50       & Yes & 0.78 & 0.77 & 0.53 \\
EfficientNet-B0 & No  & 0.80 & 0.78 & 0.52 \\
EfficientNet-B0 & Yes & 0.78 & 0.76 & 0.51 \\
ViT-B/16        & No  & 0.68 & 0.69 & 0.47 \\
ViT-B/16        & Yes & 0.69 & 0.71 & 0.49 \\
\bottomrule
\end{tabular}
\end{table}
#+end_export

** Per-Class F1 Scores

#+begin_export latex
\begin{table}[H]
\centering
\caption{Per-class F1-score for all six experiments.}
\begin{tabular}{lccccc}
\toprule
Model / CLAHE & Healthy & Mild NPDR & Moderate NPDR & Severe NPDR & Proliferative DR \\
\midrule
ResNet-50 / No        & 0.90 & 0.09 & 0.57 & 0.50 & 0.52 \\
ResNet-50 / Yes       & 0.89 & 0.17 & 0.53 & 0.45 & 0.61 \\
EfficientNet-B0 / No  & 0.90 & 0.10 & 0.57 & 0.45 & 0.58 \\
EfficientNet-B0 / Yes & 0.88 & 0.11 & 0.54 & 0.45 & 0.60 \\
ViT-B/16 / No         & 0.81 & 0.14 & 0.43 & 0.43 & 0.51 \\
ViT-B/16 / Yes        & 0.83 & 0.17 & 0.45 & 0.45 & 0.56 \\
\bottomrule
\end{tabular}
\end{table}
#+end_export

** Effect of Input Resolution (224px vs 384px)

The experiments were run twice: an initial run at 224\times{}224 and a
subsequent run at 384\times{}384 for ResNet-50 and
EfficientNet-B0. ViT-B/16 was kept at 224\times{}224 in both runs. The
resolution increase produced a substantial improvement for the CNN
models.

#+begin_export latex
\begin{table}[H]
\centering
\caption{Comparison of 224px (initial run) vs 384px (final run) test-set results for CNN models. ViT results shown for reference; both runs used 224px.}
\resizebox{\linewidth}{!}{%
\footnotesize
\begin{tabular}{llccccccc}
\toprule
Model & CLAHE & Acc (224px) & W-F1 (224px) & M-F1 (224px) & Kappa (224px) & Acc (384px) & W-F1 (384px) & M-F1 (384px) \\
\midrule
ResNet-50       & No  & 0.61 & 0.66 & 0.46 & 0.53 & 0.80 & 0.77 & 0.51 \\
ResNet-50       & Yes & 0.61 & 0.65 & 0.44 & 0.52 & 0.78 & 0.77 & 0.53 \\
EfficientNet-B0 & No  & 0.61 & 0.65 & 0.46 & 0.55 & 0.80 & 0.78 & 0.52 \\
EfficientNet-B0 & Yes & 0.62 & 0.66 & 0.45 & 0.54 & 0.78 & 0.76 & 0.51 \\
ViT-B/16        & No  & 0.68 & 0.69 & 0.47 & 0.53 & 0.68 & 0.69 & 0.47 \\
ViT-B/16        & Yes & 0.67 & 0.69 & 0.48 & 0.56 & 0.69 & 0.71 & 0.49 \\
\bottomrule
\end{tabular}%
}
\end{table}
#+end_export

Accuracy for ResNet-50 and EfficientNet-B0 improved by approximately
19 percentage points (0.61 \to 0.80) with the resolution increase. The
improvement is consistent across both CLAHE settings and is attributed
to the finer retinal microstructure (microaneurysms, hemorrhage dots,
exudates) that is present at 384px but lost at 224px.

** Confusion Matrices

*** ResNet-50 without CLAHE

#+ATTR_LATEX: :width 0.75\textwidth :placement [H]
[[file:results/resnet50_clahe=False/confusion_matrix.png]]

*** ResNet-50 with CLAHE

#+ATTR_LATEX: :width 0.75\textwidth :placement [H]
[[file:results/resnet50_clahe=True/confusion_matrix.png]]

*** EfficientNet-B0 without CLAHE

#+ATTR_LATEX: :width 0.75\textwidth :placement [H]
[[file:results/efficientnet_b0_clahe=False/confusion_matrix.png]]

*** EfficientNet-B0 with CLAHE

#+ATTR_LATEX: :width 0.75\textwidth :placement [H]
[[file:results/efficientnet_b0_clahe=True/confusion_matrix.png]]

*** ViT-B/16 without CLAHE

#+ATTR_LATEX: :width 0.75\textwidth :placement [H]
[[file:results/vit_b_16_clahe=False/confusion_matrix.png]]

*** ViT-B/16 with CLAHE

#+ATTR_LATEX: :width 0.75\textwidth :placement [H]
[[file:results/vit_b_16_clahe=True/confusion_matrix.png]]

** Training Curves

*** ResNet-50

#+ATTR_LATEX: :width \textwidth :placement [H]
[[file:results/resnet50_clahe=False/training_curves.png]]

#+ATTR_LATEX: :width \textwidth :placement [H]
[[file:results/resnet50_clahe=True/training_curves.png]]

*** EfficientNet-B0

#+ATTR_LATEX: :width \textwidth :placement [H]
[[file:results/efficientnet_b0_clahe=False/training_curves.png]]

#+ATTR_LATEX: :width \textwidth :placement [H]
[[file:results/efficientnet_b0_clahe=True/training_curves.png]]

*** ViT-B/16

#+ATTR_LATEX: :width \textwidth :placement [H]
[[file:results/vit_b_16_clahe=False/training_curves.png]]

#+ATTR_LATEX: :width \textwidth :placement [H]
[[file:results/vit_b_16_clahe=True/training_curves.png]]

* Discussion

** Resolution Is the Dominant Factor for CNN Models

The largest improvement was the ~19 percentage point accuracy jump for
ResNet-50 and EfficientNet-B0 when input resolution was increased from
224\times{}224 to 384\times{}384. Retinal images contain small
features used for diagnosis (microaneurysms, dot-like red lesions as
small as 10--20\mu{}m, dot hemorrhages, and hard exudates) that take
up only a few pixels at 224px. At 384px these features are resolved
clearly enough for the network to learn discriminative filters for
them. The resolution improvement had no effect on ViT-B/16 because
PyTorch's ViT-B/16 implementation only supports 224\times{}224 input
and cannot be run at higher resolutions without custom modifications
to the model.

** CLAHE Trades Accuracy for Minority-Class Recall

Applying CLAHE consistently reduced overall accuracy by approximately
2 percentage points for the CNN models (e.g., ResNet-50: 0.80 \to
0.78) while improving detection of minority classes. For ResNet-50,
CLAHE nearly doubled the Mild NPDR F1-score from 0.09 to 0.17, and
improved Proliferative DR F1 from 0.52 to 0.61. This happens because
CLAHE enhances the local contrast of subtle lesions, making
minority-class features more distinguishable, but the enhanced
contrast can also introduce artifacts that disrupt features the model
relied on for the majority Healthy class.

For ViT-B/16, CLAHE produced a small but consistent improvement across
most classes, and the Mild NPDR recall improved from 0.17 to
0.22. This suggests that CLAHE is more beneficial for ViT, possibly
because the transformer's attention mechanism can exploit the enhanced
contrast across longer-range spatial dependencies.

Whether CLAHE is desirable depends on the application: maximizing
overall accuracy favors CLAHE-off, while maximizing detection of the
higher-risk advanced DR grades favors CLAHE-on. For a clinical
screening tool the latter is generally preferable.

** Mild NPDR Remains Universally Difficult

Mild NPDR (class 1) had the lowest F1-score in every experiment,
ranging from 0.09 to 0.17. This class is defined by only
microaneurysms with no other lesions, a very subtle change from the
healthy retina. The confusion matrices confirm that the vast majority
of Mild NPDR predictions are misclassified as Healthy. The visual
difference is small, the class is heavily underrepresented (6.9% of
the dataset), and the class boundaries are subjective even for trained
graders. Despite weighted sampling and Focal Loss, the network does
not reliably learn to distinguish these cases. A possible solution to
this is to train with a classification model that supports higher
resolutions such as the modified Hi-ResNet, where I hypothesize that
the small lesions in class 1 get obscured or even essentially erased
when downscaling the images during the dataset preparation for use as
input to the models.

** ViT-B/16 Underperforms at 224px

ViT-B/16 achieved lower accuracy (0.68--0.69) compared to ResNet-50
and EfficientNet-B0 (0.78--0.80). This is mainly due to the resolution
constraint: ViT-B/16 uses 16\times{}16 pixel patches, so at 224px each
patch covers a 16\times{}16 area, which is large enough to subsume
entire microaneurysms. The training curves reveal severe overfitting:
the training loss approaches zero while validation loss increases
after ~10--15 epochs.  ViT-B/16 likely requires either a higher input
resolution (which demands interpolated positional embeddings and
careful fine-tuning), a larger dataset, or stronger regularization to
generalize well on medical images.

* Conclusion

Pretrained CNN architectures achieved strong diabetic retinopathy
grading performance (accuracy ~0.80, weighted F1 ~0.77--0.78) when
fine-tuned at sufficient resolution (384\times{}384). Input resolution
was the most impactful factor, with a ~19 percentage point accuracy
improvement for ResNet-50 and EfficientNet-B0 when resolution was
increased from 224 to 384 pixels. CLAHE preprocessing improved
minority class detection at a small overall accuracy cost and is
recommended for clinical settings where detecting advanced DR grades
is the priority. EfficientNet-B0 achieved the best weighted F1 (0.78)
without CLAHE, while ResNet-50 with CLAHE achieved the best macro F1
(0.53), indicating slightly more balanced performance across
classes. ViT-B/16 was constrained by its fixed-resolution patch
embedding and underperformed the CNN models.

Mild NPDR remains the hardest class: its visual distinction from a
healthy retina is subtle and the class is severely
underrepresented. Future work could address this with: dedicated data
augmentation for lesion synthesis, higher-resolution ViT variants
(ViT-L with interpolated positional embeddings), pre-training on
larger retinal image datasets, or model ensembles.

* Appendix
The repository will be available at
[[https://git.vineetk.net/eel4759_finalproject_classification/]].

#+print_bibliography: