BEGIN:VCALENDAR
VERSION:2.0
X-WR-CALNAME:EventsCalendar
PRODID:-//hacksw/handcal//NONSGML v1.0//EN
CALSCALE:GREGORIAN
BEGIN:VTIMEZONE
TZID:America/New_York
LAST-MODIFIED:20240422T053451Z
TZURL:https://www.tzurl.org/zoneinfo-outlook/America/New_York
X-LIC-LOCATION:America/New_York
BEGIN:DAYLIGHT
TZNAME:EDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0400
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZNAME:EST
TZOFFSETFROM:-0400
TZOFFSETTO:-0500
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
CATEGORIES:Uncategorized
DESCRIPTION:Advisor: Dr. Asif Kamal TurzoCommittee Members: Dr. Haiping Xu 
 and Dr. Adnan El-Nasan Abstract: Prior deep-learning vulnerability detecto
 rs report F1 scores as high as 99% on benchmark datasets, yet collapse to 
 the trivial all-safe classifier when evaluated on realistically-distribute
 d code. On Real-Vul (Chakraborty et al., 2024), a corpus that simulates de
 ployment-scale class imbalance of roughly 0.3% positive, every published m
 odel studied scores ROC-AUC near 0.50 and vulnerable-class F1 near zero. T
 his thesis demonstrates that the collapse is not fundamental. We rebuild R
 eal-Vul through a controlled Code Property Graph (CPG) pipeline, measure a
 nd remove a 37% content leak intrinsic to whole-codebase sampling, and tra
 in a relational graph neural network with a disciplined class-imbalance re
 cipe: focal loss, class-aware undersampling, and fine-tuning of a GraphCod
 eBERT node encoder. On the deduplicated test set (n = 30,873; 518 vulnerab
 le; 1.68% positive), the model reaches ROC-AUC 0.975 (95% CI [0.967, 0.983
 ]) and vulnerable-class F1 0.773 (95% CI [0.746, 0.799]), where the strong
 est prior result on the same corpus reached F1 0.46 / AUC 0.82. Three cont
 rolled experiments locate the cause and the boundary of the result. An arc
 hitectural ablation attributes the gain to the training recipe operating o
 n a protocol-matched corpus; swapping the readout architecture leaves the 
 metrics unchanged. A four-corpus transfer matrix shows that the result hol
 ds only within a single labeling protocol. A Java imbalance-reproduction e
 xperiment shows the recipe consistently beating no-handling on the same co
 rpus while reaching a strong detector only where the corpus carries learna
 ble structural signal at scale. The result is a working detector on realis
 tic data, with an account of the conditions under which it works. All CIS 
 graduate students are encouraged to attend. For further information please
  contact Dr. Asif Kamal Turzo at aturzo@umassd.edu\nEvent page: https://ww
 w.umassd.edu/events/cms/8-27-26-cis-masters-thesis-defense-by-brian-kade-b
 etterton.php
X-ALT-DESC;FMTTYPE=text/html:<html><body><p>Advisor: Dr. Asif Kamal Turzo<b
 r />Committee Members: Dr. Haiping Xu and Dr. Adnan El-Nasan</p>\n<p>Abstr
 act: Prior deep-learning vulnerability detectors report F1 scores as high 
 as 99% on benchmark datasets\, yet collapse to the trivial all-safe classi
 fier when evaluated on realistically-distributed code. On Real-Vul (Chakra
 borty et al.\, 2024)\, a corpus that simulates deployment-scale class imba
 lance of roughly 0.3% positive\, every published model studied scores ROC-
 AUC near 0.50 and vulnerable-class F1 near zero. This thesis demonstrates 
 that the collapse is not fundamental. We rebuild Real-Vul through a contro
 lled Code Property Graph (CPG) pipeline\, measure and remove a 37% content
  leak intrinsic to whole-codebase sampling\, and train a relational graph 
 neural network with a disciplined class-imbalance recipe: focal loss\, cla
 ss-aware undersampling\, and fine-tuning of a GraphCodeBERT node encoder. 
 On the deduplicated test set (n = 30\,873\; 518 vulnerable\; 1.68% positiv
 e)\, the model reaches ROC-AUC 0.975 (95% CI [0.967\, 0.983]) and vulnerab
 le-class F1 0.773 (95% CI [0.746\, 0.799])\, where the strongest prior res
 ult on the same corpus reached F1 0.46 / AUC 0.82. Three controlled experi
 ments locate the cause and the boundary of the result. An architectural ab
 lation attributes the gain to the training recipe operating on a protocol-
 matched corpus\; swapping the readout architecture leaves the metrics unch
 anged. A four-corpus transfer matrix shows that the result holds only with
 in a single labeling protocol. A Java imbalance-reproduction experiment sh
 ows the recipe consistently beating no-handling on the same corpus while r
 eaching a strong detector only where the corpus carries learnable structur
 al signal at scale. The result is a working detector on realistic data\, w
 ith an account of the conditions under which it works.</p>\n<p>All CIS gra
 duate students are encouraged to attend. For further information please co
 ntact Dr. Asif Kamal Turzo at aturzo@umassd.edu</p><p>Event page: <a href=
 "https://www.umassd.edu/events/cms/8-27-26-cis-masters-thesis-defense-by-b
 rian-kade-betterton.php">https://www.umassd.edu/events/cms/8-27-26-cis-mas
 ters-thesis-defense-by-brian-kade-betterton.php</a></a></p></body></html>
DTSTAMP:20260813T160932
DTSTART;TZID=America/New_York:20260827T140000
DTEND;TZID=America/New_York:20260827T150000
LOCATION:Dion 303
SUMMARY;LANGUAGE=en-us:Recovering Vulnerability Detection on Realistically-
 Distributed Code: Code Property Graphs, Graph Neural Networks, and a Disci
 plined Imbalance Recipe
UID:ab394d0864d38cd5b173354448c58f58@www.umassd.edu
END:VEVENT
END:VCALENDAR
