2026Journal
ASTRA: Adaptive Spatio-Temporal Representation for Action Recognition via Importance-Guided Motion Learning
J. Kim, S. Park, and H. Park
IEEE Journal of Selected Areas in Sensors — accepted, to appear — 게재 확정
Abstract초록
3D-CNN action recognizers learn appearance and motion in an entangled way, so they
lean on static visual cues instead of true motion — which hurts generalization and
interpretability on resource-constrained AIoT edge devices. ASTRA is a dual-pathway
architecture that separates semantic appearance from pure motion through three
components: an STFormer module that models space and time adaptively through
cross-attention and estimates frame-wise importance; Importance-Guided Temporal
Perturbation, which isolates appearance-free motion by comparing original and
perturbed inputs; and a Motion Gated Unit that refines the motion features through
dual-gate modulation. Under a leak-free, video-level protocol ASTRA reaches 97.88% on
UCF-101 and 81.05% on HMDB-51 with an MViT-v2-S backbone. Against its own R(2+1)D-18
backbone, the motion pathway adds a significant +2.81%p on HMDB-51 (McNemar
p < 0.001) while staying equivalent on appearance-saturated UCF-101, at 32.9M
parameters, 41.9 GFLOPs and 8.2 ms GPU latency.
UCF-101 97.88%
HMDB-51 81.05%
HMDB-51 +2.81%p · p < 0.001
32.9M params · 8.2 ms
2024Journal
Dynamic Transmission and Delay Optimization Random Access for Reduced Power Consumption
J. Kim, Y. Kim, S. Park, and H. Park
IEEE Access, vol. 12, pp. 55033–55050, April 2024
Abstract초록
Massive machine-type communications congest the random-access channel and burn power
on repeated preamble transmissions. DTDO-RA adjusts the backoff indicator (BI)
according to the number of transmissions, and uses reinforcement learning — Q-learning
and DDPG — to tune the BI together with the maximum number of preamble transmissions
(Max Tx) as the network state changes. Widening the backoff range under heavy traffic
keeps the procedure succeeding: the success rate, which falls at 10,000 UEs under the
standard procedure, holds until 15,000 UEs. At 15,000 UEs a larger BI cuts transmit
power by about 98% (3.71 to 0.072 mJ) at the cost of longer access delay, while
letting Max Tx widen the backoff range lowers the delay by up to 4.3× relative to BI
adjustment. Simulations run to 50,000 UEs and also show that shrinking Max Tx too far
raises power rather than saving it.
tx power −98.06% at 15k UEs
success-rate knee 10k → 15k UEs
4.3× lower delay than BI tuning
Q-learning · DDPG · 50k UEs
2023Journal
Limited Discriminator GAN using explainable AI model for overfitting problem
J. Kim and H. Park
ICT Express, vol. 9, no. 2, pp. 241–246, 2023
Abstract초록
GANs are widely used to augment scarce data, but a discriminator that leans too hard on
the training set overfits, and the generator ends up reproducing images that look like
the training images — augmentation that adds nothing. LDGAN opens the discriminator
with LIME to show which image regions drive its real/fake decision, and finds that an
unconstrained discriminator relies on the whole image, background included. It then
limits the discriminator's training with a threshold, set either on the discriminator's
mean loss (about 0.5) or on the generator's loss, so that the decision rests on the
object rather than the background and the generated images stay diverse. Of 500 images
generated per model, 254 from DCGAN show usable eye and nose regions, against 324 and
360 under the two LDGAN limits.
eye/nose regions in 360/500 (DCGAN 254)
D-loss limit 324/500
D-loss threshold 0.5
LIME · CIFAR-10 · Dogs vs. Cats